67.5%, measured and reproducible.
67.5% on BEAM 1M, averaged over 5 independent runs (σ 0.22%, none cherry-picked), graded by gpt-4.1-mini with the benchmark's own judging prompt.
On the same benchmark, scores vary widely with the grading protocol - the judge, the prompt, and what the retrieval layer is allowed to do. We grade with the judge the authors' own repository ships, use their judging prompt as written, change nothing to fit the test, and publish the harness, the scoring code and the reader prompt so anyone can reproduce the 67.5% exactly.
The protocol, in one table
| Item | What we do |
|---|---|
| Dataset | BEAM 1M - 35 conversations, 74,630 turns, 2.2 million memories, 700 questions |
| Judge | gpt-4.1-mini at temperature 0, running BEAM's own judging prompt - the default in the authors' repository, not a judge we picked |
| Runs | 5 independent runs, every score published, mean reported (σ 0.22%) |
| Reader | Fixed reader model and prompt, published verbatim |
What keeps it honest: a third-party judge, the reader prompt published unchanged, and every run reported - not just the best. The retrieval engine is deterministic - run it again and you get the same memories.
We climb harder benchmarks
We test on the hardest standard benchmark we haven't yet conquered - and the number is the high-water mark across every WOS model, rewritten each time a better one ships. Clear 94%, and we graduate to a harder benchmark.