How we benchmark.
On a memory benchmark, the same system can post very different numbers depending on the protocol - who grades the answers, how the prompt is written, and what the retrieval layer is allowed to do. A score means something only when you can see and reproduce the protocol behind it. That's what we publish.
As of 2 October 2026, we report three numbers for every run. The official score is graded with the benchmark's own judge and rubric. Where a benchmark names no judge, we use the same judge model our competitors used, so the numbers sit side by side. Recall is the share of the benchmark's evidence turns that WOS handed to the reader, shown as a percentage. Coverage is decided by a judge model that reads only the memories WOS returned and marks, for each question, whether they hold what the reference answer needs. Recall and coverage do not depend on the reader model; the official score does. Each benchmark is run once, and the scoring scripts are public so you can check our work.
Everything else is allowed to move, and we say so in each model's report. The benchmark dataset itself improves over time; we use the strongest public test of long-term memory available at the time of the run, and each report names the one it used. The reader model, which takes WOS's recalled memories and writes the answer, also changes as newer models arrive.
We deliberately pair WOS with the strongest reader model available, and the reason matters. A weak reader muddies the result: when an answer is wrong, you cannot tell whether WOS failed to recall the right memory, or the model simply failed to use it. A strong reader removes that ambiguity. If the right memory is in front of it, a capable model will use it, so the score reflects how well WOS retrieved rather than how well the reader coped. The better the model, the more clearly the quality of the memory shows through. The judge and the reader are always different models. The reader is the strongest model Claude and GPT each offer at the time of the run, and each report names the reader it used.
The mechanics are simple. We load each conversation history into WOS, then for every question ask WOS to recall only the relevant memories, a bounded set no matter how large the history: around 1,200 tokens on Tablet, around 3,700 on Scroll. The reader answers from those memories alone, and the judge compares the answer to the reference and scores it. Because retrieval is deterministic, the only thing that varies between runs is the reader.
A score here measures one thing: how well WOS surfaced the right memories for a reader to use. It is not a measure of your finished product, since your own prompt and model still shape the final answer, but it is comparable across languages, because the engine runs identically whatever language your users write in.
The models
Tablet is the lowest-cost way to give an agent long-term memory. Scroll reads wider and hands the reader a fuller context, for questions where completeness matters more than cost. Each carries its own report.
Every model's runs are on its own page, linked above. The full reports, including what works against us, are in research.