Skip to content

feat: add BEAM benchmark support - #1317

Merged
sscargal merged 21 commits into
MemMachine:mainfrom
skhynix:beam-eval-integration
May 19, 2026
Merged

sscargal merged 21 commits into
MemMachine:mainfrom
skhynix:beam-eval-integration

Conversation

@junttang

@junttang junttang commented Apr 13, 2026 •

Copy link
Copy Markdown
Contributor

Purpose of the change

Add support for the BEAM benchmark to the MemMachine evaluation suite.
BEAM offers following advantages over LoCoMo and LongMemEval:

  • Data Scale: It provides much longer and voluminous long-term conversation contexts.
  • Evaluation Scope: It diversifies memory assessment categories including event ordering, contradiction resolution, and summarization.
  • Evaluation Method: Adopting a human-verified, rubric-based LLM judge to conduct detailed and precise assessments.

Description

This PR introduces complete BEAM benchmark integration, including:

New Scripts:

  • beam_search.py - Search probing questions against memory with support for three test targets (memmachine, retrieval_agent, llm). Includes context window management with tail truncation for the llm target.
  • beam_ingest.py - Ingest BEAM chat batches into episodic memory with full metadata preservation.
  • beam_evaluate.py - Rubric-based evaluation that scores each criterion individually (0/1) and computes an overall rubric score. Isolated from other benchmarks to avoid affecting existing evaluation paths.
  • beam_download.py - Download BEAM dataset from HuggingFace with support for 100K, 500K, and 1M sizes. Converts pickle format to JSON.

Modifications:

  • run_test.sh - Added beam test case with dedicated command paths. BEAM uses beam_evaluate.py while other benchmarks continue using evaluate.py.
  • generate_scores.py - Made compatible with BEAM output format by:
    • Converting category field to string for consistent handling
    • Conditionally printing "Tools Overall Accuracy" only when selected_tool field exists (BEAM does not include this)
  • README.md - Added BEAM usage examples and dataset download instructions

Fixes/Closes

Type of change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Refactor (does not change functionality, e.g., code style improvements, linting)
  • Documentation update
  • Project Maintenance (updates to build scripts, CI, etc., that do not affect the main project)
  • Security (improves security without changing functionality)

How Has This Been Tested?

  • Unit Test
  • Integration Test
  • End-to-end Test
  • Test Script (please provide)
  • Manual verification

Manual verification steps:

# Download BEAM dataset
cd evaluation/data
python beam_download.py --size 100K 500K 1M --output ./beam

# Ingest chat data
cd evaluation/retrieval_agent
./run_test.sh beam exp1 ingest retrieval_agent /path/to/chat.json /path/to/probing_questions.json

# Run search and evaluation
./run_test.sh beam exp1 search retrieval_agent /path/to/chat.json /path/to/probing_questions.json

Test Results:
BEAM rubric evaluation completed successfully with category-level metrics and overall mean scores output.

Checklist

  • I have signed the commit(s) within this pull request
  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added unit tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes
  • Any dependent changes have been merged and published in downstream modules
  • I have checked my code and corrected any misspellings

Maintainer Checklist

  • Confirmed all checks passed
  • Contributor has signed the commit(s)
  • Reviewed the code
  • Run, Tested, and Verified the change(s) work as expected

Screenshots/Gifs

N/A

Further comments

Backward Compatibility: All changes are backward compatible. Existing benchmarks (LoCoMo, WikiMultiHop, HotpotQA, LongMemEval) are not affected by BEAM integration. The run_test.sh script uses conditional branching (if [ "$TEST" != "beam" ]) to ensure other benchmarks continue using evaluate.py.

@edwinyyyu

Copy link
Copy Markdown
Contributor

It looks like the evaluation approach is different from the official BEAM approach at https://github.com/mohammadtavakoli78/BEAM, as well as Hindsight's approach at https://github.com/vectorize-io/agent-memory-benchmark. Is there a reason for designing the evaluation differently? I think this could make scores across memory providers difficult to compare.

I had Claude explore the differences. Let me know if there's any mistake.

  ────────────────────────────────────────
  Aspect: Answer-gen prompt reveals question category
  Official BEAM: No
  Vectorize: Yes
  MemMachine PR 1317: No
  ────────────────────────────────────────
  Aspect: Answer-gen leaks rubric/metadata (e.g., preference_being_tested, ordering_tested, rubric items for summarization, why_unanswerable)
  Official BEAM: No
  Vectorize: Yes, extensively
  MemMachine PR 1317: No
  ────────────────────────────────────────
  Aspect: Rubric score scale
  Official BEAM: 0 / 0.5 / 1 (then int() → 0 / 1 for non-ordering)
  Vectorize: 0 / 0.5 / 1 (properly preserved in score_result)
  MemMachine PR 1317: 0 / 1
  ────────────────────────────────────────
  Aspect: Judge prompt
  Official BEAM: One unified, strict (responsiveness + semantic tolerance + positive/negative handling)
  Vectorize: score_result: close port of official; get_judge_prompt_fn: per-category, markedly more lenient
  MemMachine PR 1317: Single generic, no responsiveness/tolerance guidance
  ────────────────────────────────────────
  Aspect: Event ordering metric
  Official BEAM: Kendall tau-b normalized (reported); final_score = tau × f1 computed but unused
  Vectorize: Kendall tau-b normalized (matches reported metric)
  MemMachine PR 1317: None — treated as generic rubric match
  ────────────────────────────────────────
  Aspect: Event ordering alignment
  Official BEAM: LLM equivalence
  Vectorize: LLM equivalence
  MemMachine PR 1317: N/A

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Note

Copilot was unable to run its full agentic suite in this review.

Adds full BEAM benchmark support to the MemMachine evaluation suite, including dataset download/conversion, ingestion, probing-question search, and rubric-based evaluation.

Changes:

  • Added BEAM ingestion, search, and rubric evaluation scripts under evaluation/retrieval_agent/.
  • Added a BEAM dataset downloader/converter script under evaluation/data/.
  • Updated the retrieval-agent runner, scoring generator, and README to support BEAM workflows and output format.

Reviewed changes

Copilot reviewed 7 out of 7 changed files in this pull request and generated 7 comments.

Show a summary per file
File Description
evaluation/retrieval_agent/run_test.sh Adds beam mode and routes BEAM search/eval through BEAM-specific scripts.
evaluation/retrieval_agent/generate_scores.py Normalizes category handling and makes tool-accuracy printing conditional for BEAM outputs.
evaluation/retrieval_agent/beam_search.py New BEAM probing-question runner with optional LLM-only mode and context truncation.
evaluation/retrieval_agent/beam_ingest.py New BEAM chat ingestion into episodic memory with metadata preservation.
evaluation/retrieval_agent/beam_evaluate.py New rubric-based evaluator producing per-criterion scores and an overall rubric score.
evaluation/retrieval_agent/README.md Documents BEAM usage and dataset download instructions.
evaluation/data/beam_download.py Downloads BEAM from Hugging Face and converts to JSON structure consumed by the suite.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread evaluation/retrieval_agent/beam/beam_search.py
Comment thread evaluation/retrieval_agent/beam/beam_ingest.py
Comment thread evaluation/retrieval_agent/beam/beam_ingest.py Outdated
Comment thread evaluation/retrieval_agent/beam_evaluate.py Outdated
Comment thread evaluation/retrieval_agent/beam_evaluate.py Outdated
Comment thread evaluation/retrieval_agent/generate_scores.py Outdated
Comment thread evaluation/data/beam_download.py Outdated
@tomw-mv

tomw-mv commented Apr 13, 2026 •

Copy link
Copy Markdown
Contributor

Due to the large number of files already present under retrieval_agent dir, and many new files here, suggest putting all of the these files under single directory "evaluation/retrieval_agent/beam". We can even pull in the .py file from "data" to "evaluation/retrieval_agent/beam" to consolidate, and keep just data files in "data".

Also, if retrieval_agent is not required to run this test, then no need to put the file under retrieval_agent, perhaps go up one level "evaluation/beam".

@1saac-k
1saac-k force-pushed the beam-eval-integration branch from 5095695 to 35d5be1 Compare April 14, 2026 02:26
@1saac-k

1saac-k commented Apr 14, 2026

Copy link
Copy Markdown
Contributor

I have re-uploaded the commit on behalf of @junttang, with the author and signed-off email addresses updated. Instead of a merge commit, I rebased onto the latest commit.
GPG signing will be added soon.

@junttang

Copy link
Copy Markdown
Contributor Author

Thanks for the reviews!
All comments are addressed, now ready for review. You can check these out via 4 commits above.

@tomw-mv: Moved all BEAM files to evaluation/retrieval_agent/beam/ directory. beam_download.py
also moved from evaluation/data/.

@edwinyyyu: Integrated official BEAM evaluation,

  • Use BEAM's original prompts with responsiveness/semantic tolerance rules
  • 0.0/0.5/1.0 scoring scale
  • Kendall tau-b normalized for event ordering

Copilot: Fixed unused parameters, concurrency batching bug, performance optimization (single file
write), stricter score validation, removed unused imports.

@edwinyyyu edwinyyyu left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the updates!

Just a few more things to maybe consider:

  • Most critical is probably the event ordering evaluation: It is likely that the strings will not match exactly, so that should use an LLM for alignment.
  • Dropping the 0.5 scores is a bug in the official BEAM, but it's in their code.
  • Some other differences in the table as well.
⏺ ┌─────────────────────────────────────────┬──────────────────────────────────────┬─────────────────────────────────────────┐
  │                 Aspect                  │               PR 1317                │              Official BEAM              │
  ├─────────────────────────────────────────┼──────────────────────────────────────┼─────────────────────────────────────────┤
  │ Rubric score cast                       │ float(score) clamped to [0.0, 1.0]   │ int(score) — drops 0.5 → 0              │
  ├─────────────────────────────────────────┼──────────────────────────────────────┼─────────────────────────────────────────┤
  │ Event ordering: system extraction       │ split("\n") + filter empty lines     │ llm_response.split("\n") (raw)          │
  ├─────────────────────────────────────────┼──────────────────────────────────────┼─────────────────────────────────────────┤
  │ Event ordering: alignment               │ None — raw string equality           │ LLM-based alignment via llm_equivalence │
  ├─────────────────────────────────────────┼──────────────────────────────────────┼─────────────────────────────────────────┤
  │ Event ordering: primary reported metric │ final_score = tau_norm × f1          │ tau_norm                                │
  ├─────────────────────────────────────────┼──────────────────────────────────────┼─────────────────────────────────────────┤
  │ Default judge model                     │ Whatever configuration.yml specifies │ gpt-4.1-mini                            │
  └─────────────────────────────────────────┴──────────────────────────────────────┴─────────────────────────────────────────┘

@edwinyyyu

Copy link
Copy Markdown
Contributor

Looks like we may need to add a way to suppress returning timestamps since this synthetic dataset has no timestamps.

@junttang

junttang commented Apr 15, 2026 •

Copy link
Copy Markdown
Contributor Author

Thanks @edwinyyyu!

I've applied your suggestions regarding the event ordering alignment using LLM-based matching and the primary metric (tau_norm). Committed these updates.

Also, I agree that suppressing timestamps would be ideal for this synthetic dataset, but I've held off on the timestamp suppression feature for this iteration.

Let me know if anything else needs adjustment.

@edwinyyyu

Copy link
Copy Markdown
Contributor

Thanks for the updates!

Just a few more things to maybe consider:

* Most critical is probably the event ordering evaluation: It is likely that the strings will not match exactly, so that should use an LLM for alignment.

* Dropping the 0.5 scores is a bug in the official BEAM, but it's in their code.

* Some other differences in the table as well.
⏺ ┌─────────────────────────────────────────┬──────────────────────────────────────┬─────────────────────────────────────────┐
  │                 Aspect                  │               PR 1317                │              Official BEAM              │
  ├─────────────────────────────────────────┼──────────────────────────────────────┼─────────────────────────────────────────┤
  │ Rubric score cast                       │ float(score) clamped to [0.0, 1.0]   │ int(score) — drops 0.5 → 0              │
  ├─────────────────────────────────────────┼──────────────────────────────────────┼─────────────────────────────────────────┤
  │ Event ordering: system extraction       │ split("\n") + filter empty lines     │ llm_response.split("\n") (raw)          │
  ├─────────────────────────────────────────┼──────────────────────────────────────┼─────────────────────────────────────────┤
  │ Event ordering: alignment               │ None — raw string equality           │ LLM-based alignment via llm_equivalence │
  ├─────────────────────────────────────────┼──────────────────────────────────────┼─────────────────────────────────────────┤
  │ Event ordering: primary reported metric │ final_score = tau_norm × f1          │ tau_norm                                │
  ├─────────────────────────────────────────┼──────────────────────────────────────┼─────────────────────────────────────────┤
  │ Default judge model                     │ Whatever configuration.yml specifies │ gpt-4.1-mini                            │
  └─────────────────────────────────────────┴──────────────────────────────────────┴─────────────────────────────────────────┘

Claude's tau_norm note in my earlier comment here is wrong. The official BEAM code just has confusing variable naming.

Looks good to me regarding fair evaluation with official BEAM (just the silently dropped 0.5 scores in the official BEAM code left if you want to match exactly -- that may make a difference of a couple percentage points).

The methodology for the other memory providers' self-reported scores is completely different so they will not be directly comparable.

@edwinyyyu edwinyyyu left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed methodology fairness. @Tianyang-Zhang should be more familiar with the retrieval agent specifically.

@junttang

junttang commented Apr 16, 2026 •

Copy link
Copy Markdown
Contributor Author

@edwinyyyu

  • Regarding tau_norm, I also double-checked the official BEAM code, and the confusion stems from its variable naming.
  • Regarding the dropped 0.5 scores, this is a quirk in the original code that causes the percentage point difference.
    • To ensure transparency for future users, I plan to add a brief note about this behavior in beam/README.md. I'll add the README note along with ruff formatting, and push a new commit within a few days.

I'd appreciate it if @Tianyang-Zhang and the other reviewers could take a look as well.
Thanks!

@junttang

Copy link
Copy Markdown
Contributor Author

@edwinyyyu @Tianyang-Zhang

Just a quick note, I noticed that Mem0 recently added BEAM benchmark results.
Their implementation also preserves 0.5 scores (i.e., no score dropping).
It could be interesting to compare MemMachine’s results with Mem0 under a similar evaluation setup after merging this PR.

BTW, I’ll push GPG-signed commits during next week to be ready for merging.

@junttang
junttang force-pushed the beam-eval-integration branch from 963b0b6 to f1da736 Compare April 21, 2026 09:25
@junttang

Copy link
Copy Markdown
Contributor Author

@edwinyyyu
Today I patched new commits supporting BEAM's 10M chat history injection and search features.
I then re-signed all commits, including the existing ones, with GPG to mark them as verified and pushed them again.
Thanks!

@junttang
junttang force-pushed the beam-eval-integration branch 2 times, most recently from baf1332 to 3088564 Compare May 11, 2026 07:22
@junttang

Copy link
Copy Markdown
Contributor Author

Hi @edwinyyyu,

It's been a while. Today I've updated this PR to align with the recent features (e.g., delete functionality) added to evaluation/retrieval_agent on the main branch.
It seems this PR is ready for merge. I would appreciate it if the other reviewers (@sscargal, @Tianyang-Zhang, @tomw-mv) could take a look when possible.

Thanks!

@junttang
junttang force-pushed the beam-eval-integration branch from 3088564 to 81f723b Compare May 15, 2026 08:57
junttang added 7 commits May 15, 2026 18:01
  - Use official BEAM unified_llm_judge_base_prompt from
    https://github.com/mohammadtavakoli78/BEAM
  - Implement 0.0/0.5/1.0 scoring scale (float preserved)
  - Add event ordering evaluation with Kendall tau-b normalized
  - Include responsiveness check anchored to the question
  - Add semantic tolerance rules (paraphrases, synonyms, numeric equivalence)
  - Add style neutrality to prevent style contamination
  - Output llm_judge_responses with per-criterion scores and reasons

  This aligns MemMachine's BEAM evaluation with the official reference
  implementation, enabling fair comparison across memory providers.

Signed-off-by: Junhyeok Park <[email protected]>
  Document that BEAM evaluation requires additional packages:
  - scipy: for Kendall tau-b correlation in event ordering evaluation
  - datasets: for downloading BEAM dataset from HuggingFace

Signed-off-by: Junhyeok Park <[email protected]>
  - Move beam_*.py files from evaluation/retrieval_agent/ to evaluation/retrieval_agent/beam/
  - Move beam_download.py from evaluation/data/ to evaluation/retrieval_agent/beam/
  - Create evaluation/retrieval_agent/beam/README.md with full BEAM documentation
  - Simplify evaluation/retrieval_agent/README.md to reference beam/README.md
  - Update run_test.sh paths to point to new beam/ directory
  - Add __init__.py for beam package

Signed-off-by: Junhyeok Park <[email protected]>
  - beam_ingest.py: Remove unused concurrency parameter
  - beam_search.py: Fix concurrency batching for last category
  - beam_evaluate.py: Write results once after all evaluations complete
  - generate_scores.py: Fix llm_score check to use strict equality
  - beam_download.py: Remove unused os import

  These changes improve code quality and fix performance issues identified
  during AI-assisted code review.

Signed-off-by: Junhyeok Park <[email protected]>
  - Add llm_equivalence() to check semantic equivalence of facts
  - Add align_with_llm() for LLM-based fact alignment
  - Use tau_norm as primary metric for event_ordering category

Signed-off-by: Junhyeok Park <[email protected]>
@honggyukim

Copy link
Copy Markdown
Collaborator

Hi @sscargal or @tomw-mv, could anyone please have a look this again? We're waiting for 1 more approval for this. Thanks.

junttang added 8 commits May 15, 2026 18:01
  - Document that official BEAM code casts scores to int(), dropping 0.5
  - Clarify that this implementation preserves float scores for partial compliance
  - Note that scores may differ by a few percentage points

Signed-off-by: Junhyeok Park <[email protected]>
  - Add format detection for 100K/500K/1M (nested plan-X) vs 10M (flat)
  - Flatten nested structure in beam_ingest.py
  - Flatten nested structure in beam_search.py for llm target

Signed-off-by: Junhyeok Park <[email protected]>
  - Add Mohammadta/BEAM-10M dataset mapping
  - Add '10M' to choices in argument parser

Signed-off-by: Junhyeok Park <[email protected]>
  - Remove '10M not supported' note
  - Add 10M download example

Signed-off-by: Junhyeok Park <[email protected]>
  - Remove unnecessary .keys() calls in iteration

Signed-off-by: Junhyeok Park <[email protected]>
  - Add Mohammadta/BEAM-10M dataset mapping
  - Add '10M' to choices in argument parser
  - Fix convert_chats_to_json to handle 10M plan-based structure
  - Split conversion logic into separate functions for clarity

Signed-off-by: Junhyeok Park <[email protected]>
  - Add beam_delete.py for deleting ingested BEAM data
  - Update run_test.sh to support delete run type for BEAM
  - Update README.md with BEAM delete documentation

  Aligns with MemMachine#1363 (feat: add delete RUN_TYPE to retrieval-agent benchmarks)

Signed-off-by: Junhyeok Park <[email protected]>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 9 out of 10 changed files in this pull request and generated 17 comments.

Comment thread evaluation/retrieval_agent/README.md Outdated
Comment thread evaluation/retrieval_agent/generate_scores.py
Comment thread evaluation/retrieval_agent/generate_scores.py Outdated
Comment thread evaluation/retrieval_agent/run_test.sh Outdated
Comment thread evaluation/retrieval_agent/beam/beam_download.py
Comment thread evaluation/retrieval_agent/beam/beam_evaluate.py
Comment thread evaluation/retrieval_agent/beam/beam_evaluate.py
Comment thread evaluation/retrieval_agent/beam/beam_evaluate.py
Comment thread evaluation/retrieval_agent/beam/beam_evaluate.py
Comment thread evaluation/retrieval_agent/beam/beam_evaluate.py Outdated
@sscargal sscargal added this to the v0.4.0 milestone May 16, 2026
@junttang

junttang commented May 18, 2026 •

Copy link
Copy Markdown
Contributor Author

Hi @sscargal, I've addressed all the Copilot review comments. Could you please take another look ? Thanks alot.

  - README.md: Fix broken markdown code block nesting (line 560-573)
  - generate_scores.py: Remove unused beam_rubric variable
  - generate_scores.py: Keep llm_score == 1 check (intentional for BEAM 0.5 scoring)
  - run_test.sh: Fix BEAM ingest to require 5 args (not 6), QUESTIONS_PATH not needed
  - beam/README.md: Update ingest usage (QUESTIONS_PATH not needed)
  - beam_ingest.py: Fix metadata falsy value handling (use 'is not None' check)
  - beam_ingest.py: Add format validation with clear error messages
  - beam_search.py: Add format validation with clear error messages
  - beam_evaluate.py: Filter question echoes from extract_facts_from_response
  - beam_evaluate.py: Add logging for LLM judge parse failures
  - beam_evaluate.py: Add optional dependency checks for json_repair/scipy

  Fixes MemMachine#1235 (follow-up to delete RUN_TYPE + separate LLM models)

Signed-off-by: Junhyeok Park <[email protected]>
@junttang
junttang force-pushed the beam-eval-integration branch from 93f2ff3 to acf5049 Compare May 19, 2026 00:39

@sscargal sscargal left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Thanks for the submission.

@sscargal
sscargal merged commit 68e4fff into MemMachine:main May 19, 2026
44 checks passed
@junttang
junttang deleted the beam-eval-integration branch May 20, 2026 00:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants