Tablet 1 retrieves once and hands back what it found. Tablet 2 keeps that pass and adds a second move: when the first result does not carry the answer, the caller can ask again, and the engine returns memories it has not already delivered rather than the same ones twice, up to three times. No language model runs in that loop. On LongMemEval-S, Tablet 2 scores a mean of 95.7% across three runs, with a standard deviation of 0.4%; the individual runs were 95.2, 96.0, and 96.0. Read by a different frontier model, the same runs score 93.7% (σ 0.3), and both figures are published here for the reason given below. On BEAM 1M, 700 questions over 2.2 million memories, it scores 67.5% across five runs (σ 0.22). The two benchmarks are not on one scale and this page keeps them apart.
temperature 0 running BEAM's own judging prompt, the default in the authors' repository. Grading is a separate pass from answering. BEAM fixes neither the reader model nor its prompt, so we name ours: Opus 5, on a five-line prompt.Every configuration we ran, and the per-category scores behind each one. Nothing is summarised away. The second row is BEAM 1M, a second benchmark on the same engine build: a corpus 40× larger, ten abilities instead of six, and a judge that scores each answer against a checklist rather than as right or wrong. Its numbers are not on one scale with LongMemEval's.
| Category | Run 1 | Run 2 | Run 3 | Average |
|---|---|---|---|---|
| Single-session preference | 100.0 | 100.0 | 100.0 | 100.0 |
| Single-session assistant | 100.0 | 100.0 | 96.4 | 98.8 |
| Temporal reasoning | 97.0 | 97.0 | 97.0 | 97.0 |
| Single-session user | 95.7 | 94.3 | 95.7 | 95.2 |
| Knowledge update | 93.6 | 96.2 | 94.9 | 94.9 |
| Multi-session | 91.0 | 93.2 | 94.7 | 93.0 |
| Overall | 95.2 | 96.0 | 96.0 | 95.7 |
Mean 95.7% · standard deviation 0.4% · every run published, 0 cherry-picks · 500 / 500 graded in every run.
Tablet 2 adds one move, and it takes a sentence to state and was hard to build. A memory engine that answers once has to decide, in that one pass, everything it will ever say about a question. It has no way of knowing whether what it found was enough, because nothing tells it afterwards. So it does what any careful system does under that constraint: it returns a little more than needed, in case.
Asking again removes that guess. The reader looks at what came back, sees that the answer is not in it, and says so. What returns then is not the same handful shifted by one, but the memories the engine has been keeping track of not having given you yet. It will do that up to three times. Because the engine is the one holding that record, no language model has to sit in the path to make it work, and asking again stays as fast, as steady, and as language-neutral as the first pass.
BEAM 1M · 700 questions · 5 runs
Turning verify off is the only change: same engine, same corpus, same reader and judge. Tokens were counted with a tokenizer rather than estimated from character length. It does not pay everywhere — questions whose answer was already in the first pass lose ground, which is why verify is a setting you pass per call.
The numbers land where that reasoning predicts, and the comparison that shows it is the one where only a single setting moves. Hold the engine, the corpus, the reader and the judge still, and switch verify off: LongMemEval-S falls from 95.0% to 93.8%, and BEAM 1M from 67.5% to 58.6%. The gain concentrates where one pass is least likely to have gathered everything. Summarization on BEAM 1M needs coverage across sessions rather than one good hit, and it moves 31.4% to 52.5%. It is not free everywhere. Single-session-user questions on LongMemEval-S lose 4.2 points, because asking again when the answer was already in hand adds material that competes with it. Asking again is worth the most exactly where asking once was not enough, and worth less than nothing where it was.
LongMemEval-S asks whether a fact can be found inside a long conversation. BEAM 1M asks something else: whether memory holds its shape at a scale no context window reaches, 2.2 million memories drawn from 74,630 turns. Both are on this page, run by run, and it is the second one that shows the shape of Tablet 2 most plainly. Answering once, it scores 58.6%. Allowed to ask again, 67.5%.
Against our own earlier engines the honest comparison is narrower than the headline. Tablet and Scroll are separate tiers rather than successive generations, and the only one where we control both sides is Scroll 1.2 read by the same model: 92.3 to 93.7, a gain of 1.4 points. Those two figures come from separate campaigns rather than paired runs and their intervals overlap, so we report 1.4 as what we measured with the reader held fixed, not as an established difference. Against Tablet 1 there is no controlled comparison at all. Its 85.2% was read by a different model, and the per-question comparison we did produce was withdrawn: a generation claim would need Tablet 1 re-run on the same corpus, the same reader and the same verify budget, and we have not done that.
Tablet 2 also remembers images. A memory can be an image, and the engine indexes the image itself rather than words about it, so a sentence typed in any language finds it even when the record carries no caption at all. That is the condition worth naming, because it is the one where keyword search has nothing to match: on caption-less images Tablet 2 reaches 91.4% recall@5 and BM25 scores zero, having no text to score. The full protocol is in the paper.
Crossmodal-3600 · 300 images · recall@5
The engine indexes the image rather than words about it, so a sentence in any language finds it with no caption stored. Eleven of the fourteen languages sit above 95, and the two at the bottom are our worst numbers anywhere. BM25 is 0.0 in every one of these cells, not because it scores badly but because a record with no text gives it nothing to score.
An average would hide what matters about that. A retriever can look strong across fourteen languages while being a strong English retriever with thirteen weak ones attached, so the number to read is not the mean but the spread. Tablet 2 varies by 14.0 points across those languages; the two open baselines we could run vary by 27.5 and 27.7, and one of them scores 91.0% on English and 4.7% on Russian from the identical stored image. Avoiding keywords is not the same thing as being language-neutral, and the only way to tell the two apart is to look at the worst language rather than the average one.
Two results here work against us, and they belong on the same page as the rest. Retrieval degrades sharply in low-resource languages: Swahili at 53.0% and Telugu at 64.0% are the lowest numbers this model produces anywhere, though the strongest open baseline we could run reaches 6.0% and 5.0% on those same two. And attaching an English caption to an image makes it harder to find in every other language, by 11.4 points on average. That is a defect in something you can buy rather than an artifact of the test, because real stores hold exactly the mixture that produces it. If your users search in more than one language, store images without captions. The other half of the same measurement belongs here too: across those fourteen languages this model averages 91.4% recall@5, and on a controlled corpus of seventy store-language by query-language cells it averages 95.2% where a word-matching baseline, given the better of two tokenizers in every cell, averages 19.0% and scores exactly zero in fifty-four of the seventy. This model scores zero in none of them.
The third change is smaller and is aimed at the model rather than at you. An assistant leaning on a store it did not write has no way of knowing how far that memory has been edited underneath it, and the honest thing is to let it ask. Won is where that kind of question lives, and revisions is the first one: how many memories in this store were changed after they were written, against how many were not.
measured 2026-08-14 · median per call
An assistant leaning on a store it did not write should be able to ask how far that memory has been edited underneath it. Won is where those questions live and revisions is the first: two counts, and no model of any kind runs. It carries no charge, because putting a price on it would make the honest answer the expensive one and callers would stop asking.
It answers with counts and nothing else, which is the part that makes it usable. The response is the same size for a store of a hundred memories and a store of a hundred million, so a model can ask it in the middle of a conversation without the answer growing into the context it was trying to protect. It is free for the same reason: putting a price on it would make the honest answer the expensive one, and callers would stop asking. A store where three facts in ten have been replaced deserves less confidence than one nobody has edited, and that is a thing worth knowing before leaning on it.
Who ran this. An internal run. Every number here is self-reported, and so is every number we place it beside.
Corpus defects. 211 of the 74,630 turns were truncated by our ingest harness and one failed to store. Our fault, not the benchmark's, and both stayed in the measured corpus.