Proof

67.5%, measured and reproducible.

67.5% on BEAM 1M, averaged over 5 independent runs (σ 0.22%, none cherry-picked), graded by gpt-4.1-mini with the benchmark's own judging prompt.

On the same benchmark, scores vary widely with the grading protocol - the judge, the prompt, and what the retrieval layer is allowed to do. We grade with the judge the authors' own repository ships, use their judging prompt as written, change nothing to fit the test, and publish the harness, the scoring code and the reader prompt so anyone can reproduce the 67.5% exactly.

The protocol, in one table

ItemWhat we do
DatasetBEAM 1M - 35 conversations, 74,630 turns, 2.2 million memories, 700 questions
Judgegpt-4.1-mini at temperature 0, running BEAM's own judging prompt - the default in the authors' repository, not a judge we picked
Runs5 independent runs, every score published, mean reported (σ 0.22%)
ReaderFixed reader model and prompt, published verbatim

What keeps it honest: a third-party judge, the reader prompt published unchanged, and every run reported - not just the best. The retrieval engine is deterministic - run it again and you get the same memories.

We climb harder benchmarks

We test on the hardest standard benchmark we haven't yet conquered - and the number is the high-water mark across every WOS model, rewritten each time a better one ships. Clear 94%, and we graduate to a harder benchmark.

BEAM 1MIn progress
Tablet67.5%
gpt-4.1-mini judge · five-run mean94% to graduate
Earlier benchmark LongMemEval-S Cleared
Tablet95.7%
Scroll92.3%
GPT-4o judge · best across all WOS models94% to graduate
See the full report