Repository navigation
feat: add BEAM benchmark support - #1317
Conversation
|
It looks like the evaluation approach is different from the official BEAM approach at https://github.com/mohammadtavakoli78/BEAM, as well as Hindsight's approach at https://github.com/vectorize-io/agent-memory-benchmark. Is there a reason for designing the evaluation differently? I think this could make scores across memory providers difficult to compare. I had Claude explore the differences. Let me know if there's any mistake. |
There was a problem hiding this comment.
Pull request overview
Note
Copilot was unable to run its full agentic suite in this review.
Adds full BEAM benchmark support to the MemMachine evaluation suite, including dataset download/conversion, ingestion, probing-question search, and rubric-based evaluation.
Changes:
- Added BEAM ingestion, search, and rubric evaluation scripts under
evaluation/retrieval_agent/. - Added a BEAM dataset downloader/converter script under
evaluation/data/. - Updated the retrieval-agent runner, scoring generator, and README to support BEAM workflows and output format.
Reviewed changes
Copilot reviewed 7 out of 7 changed files in this pull request and generated 7 comments.
Show a summary per file
| File | Description |
|---|---|
| evaluation/retrieval_agent/run_test.sh | Adds beam mode and routes BEAM search/eval through BEAM-specific scripts. |
| evaluation/retrieval_agent/generate_scores.py | Normalizes category handling and makes tool-accuracy printing conditional for BEAM outputs. |
| evaluation/retrieval_agent/beam_search.py | New BEAM probing-question runner with optional LLM-only mode and context truncation. |
| evaluation/retrieval_agent/beam_ingest.py | New BEAM chat ingestion into episodic memory with metadata preservation. |
| evaluation/retrieval_agent/beam_evaluate.py | New rubric-based evaluator producing per-criterion scores and an overall rubric score. |
| evaluation/retrieval_agent/README.md | Documents BEAM usage and dataset download instructions. |
| evaluation/data/beam_download.py | Downloads BEAM from Hugging Face and converts to JSON structure consumed by the suite. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
|
Due to the large number of files already present under retrieval_agent dir, and many new files here, suggest putting all of the these files under single directory "evaluation/retrieval_agent/beam". We can even pull in the .py file from "data" to "evaluation/retrieval_agent/beam" to consolidate, and keep just data files in "data". Also, if retrieval_agent is not required to run this test, then no need to put the file under retrieval_agent, perhaps go up one level "evaluation/beam". |
5095695 to
35d5be1
Compare
|
I have re-uploaded the commit on behalf of @junttang, with the author and signed-off email addresses updated. Instead of a merge commit, I rebased onto the latest commit. |
|
Thanks for the reviews! @tomw-mv: Moved all BEAM files to @edwinyyyu: Integrated official BEAM evaluation,
Copilot: Fixed unused parameters, concurrency batching bug, performance optimization (single file |
There was a problem hiding this comment.
Thanks for the updates!
Just a few more things to maybe consider:
- Most critical is probably the event ordering evaluation: It is likely that the strings will not match exactly, so that should use an LLM for alignment.
- Dropping the 0.5 scores is a bug in the official BEAM, but it's in their code.
- Some other differences in the table as well.
⏺ ┌─────────────────────────────────────────┬──────────────────────────────────────┬─────────────────────────────────────────┐
│ Aspect │ PR 1317 │ Official BEAM │
├─────────────────────────────────────────┼──────────────────────────────────────┼─────────────────────────────────────────┤
│ Rubric score cast │ float(score) clamped to [0.0, 1.0] │ int(score) — drops 0.5 → 0 │
├─────────────────────────────────────────┼──────────────────────────────────────┼─────────────────────────────────────────┤
│ Event ordering: system extraction │ split("\n") + filter empty lines │ llm_response.split("\n") (raw) │
├─────────────────────────────────────────┼──────────────────────────────────────┼─────────────────────────────────────────┤
│ Event ordering: alignment │ None — raw string equality │ LLM-based alignment via llm_equivalence │
├─────────────────────────────────────────┼──────────────────────────────────────┼─────────────────────────────────────────┤
│ Event ordering: primary reported metric │ final_score = tau_norm × f1 │ tau_norm │
├─────────────────────────────────────────┼──────────────────────────────────────┼─────────────────────────────────────────┤
│ Default judge model │ Whatever configuration.yml specifies │ gpt-4.1-mini │
└─────────────────────────────────────────┴──────────────────────────────────────┴─────────────────────────────────────────┘
|
Looks like we may need to add a way to suppress returning timestamps since this synthetic dataset has no timestamps. |
|
Thanks @edwinyyyu! I've applied your suggestions regarding the event ordering alignment using LLM-based matching and the primary metric (tau_norm). Committed these updates. Also, I agree that suppressing timestamps would be ideal for this synthetic dataset, but I've held off on the timestamp suppression feature for this iteration. Let me know if anything else needs adjustment. |
Claude's tau_norm note in my earlier comment here is wrong. The official BEAM code just has confusing variable naming. Looks good to me regarding fair evaluation with official BEAM (just the silently dropped 0.5 scores in the official BEAM code left if you want to match exactly -- that may make a difference of a couple percentage points). The methodology for the other memory providers' self-reported scores is completely different so they will not be directly comparable. |
edwinyyyu
left a comment
There was a problem hiding this comment.
Reviewed methodology fairness. @Tianyang-Zhang should be more familiar with the retrieval agent specifically.
I'd appreciate it if @Tianyang-Zhang and the other reviewers could take a look as well. |
|
Just a quick note, I noticed that Mem0 recently added BEAM benchmark results. BTW, I’ll push GPG-signed commits during next week to be ready for merging. |
963b0b6 to
f1da736
Compare
|
@edwinyyyu |
baf1332 to
3088564
Compare
|
Hi @edwinyyyu, It's been a while. Today I've updated this PR to align with the recent features (e.g., Thanks! |
3088564 to
81f723b
Compare
Signed-off-by: Junhyeok Park <[email protected]>
Signed-off-by: Junhyeok Park <[email protected]>
Signed-off-by: Junhyeok Park <[email protected]>
Signed-off-by: Junhyeok Park <[email protected]>
Signed-off-by: Junhyeok Park <[email protected]>
Signed-off-by: Junhyeok Park <[email protected]>
Signed-off-by: Junhyeok Park <[email protected]>
- Use official BEAM unified_llm_judge_base_prompt from
https://github.com/mohammadtavakoli78/BEAM
- Implement 0.0/0.5/1.0 scoring scale (float preserved)
- Add event ordering evaluation with Kendall tau-b normalized
- Include responsiveness check anchored to the question
- Add semantic tolerance rules (paraphrases, synonyms, numeric equivalence)
- Add style neutrality to prevent style contamination
- Output llm_judge_responses with per-criterion scores and reasons
This aligns MemMachine's BEAM evaluation with the official reference
implementation, enabling fair comparison across memory providers.
Signed-off-by: Junhyeok Park <[email protected]>
Document that BEAM evaluation requires additional packages: - scipy: for Kendall tau-b correlation in event ordering evaluation - datasets: for downloading BEAM dataset from HuggingFace Signed-off-by: Junhyeok Park <[email protected]>
- Move beam_*.py files from evaluation/retrieval_agent/ to evaluation/retrieval_agent/beam/ - Move beam_download.py from evaluation/data/ to evaluation/retrieval_agent/beam/ - Create evaluation/retrieval_agent/beam/README.md with full BEAM documentation - Simplify evaluation/retrieval_agent/README.md to reference beam/README.md - Update run_test.sh paths to point to new beam/ directory - Add __init__.py for beam package Signed-off-by: Junhyeok Park <[email protected]>
- beam_ingest.py: Remove unused concurrency parameter - beam_search.py: Fix concurrency batching for last category - beam_evaluate.py: Write results once after all evaluations complete - generate_scores.py: Fix llm_score check to use strict equality - beam_download.py: Remove unused os import These changes improve code quality and fix performance issues identified during AI-assisted code review. Signed-off-by: Junhyeok Park <[email protected]>
- Add llm_equivalence() to check semantic equivalence of facts - Add align_with_llm() for LLM-based fact alignment - Use tau_norm as primary metric for event_ordering category Signed-off-by: Junhyeok Park <[email protected]>
Signed-off-by: Junhyeok Park <[email protected]>
- Document that official BEAM code casts scores to int(), dropping 0.5 - Clarify that this implementation preserves float scores for partial compliance - Note that scores may differ by a few percentage points Signed-off-by: Junhyeok Park <[email protected]>
- Add format detection for 100K/500K/1M (nested plan-X) vs 10M (flat) - Flatten nested structure in beam_ingest.py - Flatten nested structure in beam_search.py for llm target Signed-off-by: Junhyeok Park <[email protected]>
- Add Mohammadta/BEAM-10M dataset mapping - Add '10M' to choices in argument parser Signed-off-by: Junhyeok Park <[email protected]>
- Remove '10M not supported' note - Add 10M download example Signed-off-by: Junhyeok Park <[email protected]>
- Remove unnecessary .keys() calls in iteration Signed-off-by: Junhyeok Park <[email protected]>
- Add Mohammadta/BEAM-10M dataset mapping - Add '10M' to choices in argument parser - Fix convert_chats_to_json to handle 10M plan-based structure - Split conversion logic into separate functions for clarity Signed-off-by: Junhyeok Park <[email protected]>
- Add beam_delete.py for deleting ingested BEAM data - Update run_test.sh to support delete run type for BEAM - Update README.md with BEAM delete documentation Aligns with MemMachine#1363 (feat: add delete RUN_TYPE to retrieval-agent benchmarks) Signed-off-by: Junhyeok Park <[email protected]>
81f723b to
ec7ca0e
Compare
|
Hi @sscargal, I've addressed all the Copilot review comments. Could you please take another look ? Thanks alot. |
- README.md: Fix broken markdown code block nesting (line 560-573) - generate_scores.py: Remove unused beam_rubric variable - generate_scores.py: Keep llm_score == 1 check (intentional for BEAM 0.5 scoring) - run_test.sh: Fix BEAM ingest to require 5 args (not 6), QUESTIONS_PATH not needed - beam/README.md: Update ingest usage (QUESTIONS_PATH not needed) - beam_ingest.py: Fix metadata falsy value handling (use 'is not None' check) - beam_ingest.py: Add format validation with clear error messages - beam_search.py: Add format validation with clear error messages - beam_evaluate.py: Filter question echoes from extract_facts_from_response - beam_evaluate.py: Add logging for LLM judge parse failures - beam_evaluate.py: Add optional dependency checks for json_repair/scipy Fixes MemMachine#1235 (follow-up to delete RUN_TYPE + separate LLM models) Signed-off-by: Junhyeok Park <[email protected]>
93f2ff3 to
acf5049
Compare
sscargal
left a comment
There was a problem hiding this comment.
LGTM. Thanks for the submission.
Purpose of the change
Add support for the BEAM benchmark to the MemMachine evaluation suite.
BEAM offers following advantages over LoCoMo and LongMemEval:
Description
This PR introduces complete BEAM benchmark integration, including:
New Scripts:
beam_search.py- Search probing questions against memory with support for three test targets (memmachine, retrieval_agent, llm). Includes context window management with tail truncation for thellmtarget.beam_ingest.py- Ingest BEAM chat batches into episodic memory with full metadata preservation.beam_evaluate.py- Rubric-based evaluation that scores each criterion individually (0/1) and computes an overall rubric score. Isolated from other benchmarks to avoid affecting existing evaluation paths.beam_download.py- Download BEAM dataset from HuggingFace with support for 100K, 500K, and 1M sizes. Converts pickle format to JSON.Modifications:
run_test.sh- Addedbeamtest case with dedicated command paths. BEAM usesbeam_evaluate.pywhile other benchmarks continue usingevaluate.py.generate_scores.py- Made compatible with BEAM output format by:selected_toolfield exists (BEAM does not include this)README.md- Added BEAM usage examples and dataset download instructionsFixes/Closes
Type of change
How Has This Been Tested?
Manual verification steps:
Test Results:
BEAM rubric evaluation completed successfully with category-level metrics and overall mean scores output.
Checklist
Maintainer Checklist
Screenshots/Gifs
N/A
Further comments
Backward Compatibility: All changes are backward compatible. Existing benchmarks (LoCoMo, WikiMultiHop, HotpotQA, LongMemEval) are not affected by BEAM integration. The
run_test.shscript uses conditional branching (if [ "$TEST" != "beam" ]) to ensure other benchmarks continue usingevaluate.py.