Scroll 1 is the next WOS memory model after Tablet, built for the questions that a single, lean pass tends to leave half-answered. We publish it the same way we publish every model, laying out the full protocol, every run, and nothing cherry-picked. On LongMemEval-S, Scroll 1 scores a mean of 90.7% across five runs, with a standard deviation of 0.5%, a clear step up from Tablet 1's 85.2%. The individual runs were 90.4, 90.0, 91.6, 90.8, and 90.8, and all five of them landed at or above 90.
From that moment, any call that names the model scroll-1 is refused. There is one API host, and the model is chosen per request, so nothing about your endpoint or your keys changes.
Scroll 1.2 replaces Scroll 1. It costs the same, answers the same calls, and it adds the Memoir and Archive delivery forms. To move over, change the model name: send the header X-WOS-Model: scroll-1.2, or call withModel("scroll-1.2") in the SDK. Nothing else in your code changes. Tablet 2 also replaces Scroll 1, if you would rather have the newer engine at Tablet's lower price.
Every Tablet model and every Scroll model reads the same memory, so there is nothing to migrate. What you stored through Scroll 1 is the memory that Scroll 1.2 and Tablet 2 answer from.
Questions or help moving over: support@wontopos.com
temperature 0, returning a plain yes / no under the official LongMemEval per-category rules. Both prompts are published below.Same benchmark, same protocol. Where the fuller read changes the answer, and where it barely matters.
| Category | Tablet 1 | Scroll 1 | Δ |
|---|---|---|---|
| Single-session user | 98.9 | 99.2 | +0.3 |
| Single-session assistant | 96.4 | 96.8 | +0.4 |
| Knowledge update | 91.3 | 94.1 | +2.8 |
| Preference inference | 86.0 | 97.3 | +11.3 |
| Temporal reasoning | 80.7 | 87.1 | +6.4 |
| Multi-session | 73.8 | 83.9 | +10.1 |
| Overall | 85.2 | 90.7 | +5.5 |
The largest gains are in preference inference (+11.3), multi-session (+10.1), and temporal reasoning (+6.3), which are exactly the cases where one missing piece changes the answer. On single-session recall, already near the ceiling, the two stay within half a point of each other.
The raw scores behind the 90.7%.
| Category | Run 1 | Run 2 | Run 3 | Run 4 | Run 5 | Average |
|---|---|---|---|---|---|---|
| Single-session user | 98.6 | 98.6 | 98.6 | 100.0 | 100.0 | 99.2 |
| Single-session assistant | 98.2 | 96.4 | 96.4 | 98.2 | 94.6 | 96.8 |
| Knowledge update | 92.3 | 93.6 | 94.9 | 92.3 | 97.4 | 94.1 |
| Preference inference | 96.7 | 100.0 | 96.7 | 100.0 | 93.3 | 97.3 |
| Temporal reasoning | 85.7 | 84.2 | 90.2 | 86.5 | 88.7 | 87.1 |
| Multi-session | 85.0 | 84.2 | 84.2 | 84.2 | 82.0 | 83.9 |
| Overall | 90.4 | 90.0 | 91.6 | 90.8 | 90.8 | 90.7 |
Mean 90.7% · standard deviation 0.5% · range 90.0–91.6 · all 5 runs published, 0 cherry-picks.
Retrieval is a genuinely hard problem, and the field has earned real ground on it. WOS is no exception. A single query surfaces the right evidence in the top results 99.6% of the time. But finding the exact memory is only half of it. The harder half is feeding that memory to a model so it understands and answers correctly, and that half is not free. Tablet 1 reads from a single query and answers 85.2% of LongMemEval-S correctly. Scroll 1 delivers the same evidence more fully, about 3× the context, and answers 90.7%. The memory found is identical. The points between them are delivery, and that is the problem we take on.
Both models surface the evidence 99.6–100% of the time, however the query is shaped. Retrieval is saturated.
The memory is right there at 99.6%, yet the model answers 85.2% (Tablet) to 90.7% (Scroll). The distance below the dotted line is the delivery problem.
Retrieval finds the gold evidence 99.6% of the time and holds there however the query is shaped. It is saturated. The answer score does not climb that high. Tablet 1 turns the found memory into a correct answer 85.2% of the time; Scroll 1's fuller read brings that to 90.7%. The roughly nine points still sitting between 90.7% and near-perfect retrieval are the honest size of the delivery problem: the memory is right there, and a model still does not always use it. That gap, not retrieval, is where the work and the value are. (Left: session-level retrieval recall at each top-k, y-axis zoomed to the 99–100% band. Right: delivered answer accuracy against the 99.6% retrieval ceiling. Answer accuracy: five-run mean on LongMemEval-S, graded by an independent model under the official rules. Retrieval: each question isolated to its own ~53-session history.)
Curves show the shape of the distribution around the measured median; exact spread varies by workload.
Some of what you remember does not live in any one place. A person's reasons, the direction they quietly lean, the shape a plan takes as it forms over time. Things like these are never set down all at once. They are laid a little at a time, across many days and many separate conversations, and no single moment ever holds the whole of them. To answer well about something like that, one memory is not enough. You have to gather the pieces that were scattered and hold them together at the same time. Tablet brings back the one piece that best fits the question, and little else beside it. Scroll was made for the other kind of question, the kind where that single piece was never going to be the whole answer.
So Scroll reaches wider on every call. Where Tablet finds the single closest memory and stops there, Scroll keeps following the grain outward, into the turns that were sitting next to it and the moments that stood beside the one it found, and it brings back a fuller stretch of context, roughly three times as much. The effect is that the scattered pieces arrive together, gathered into one hand, instead of arriving one short of a whole. This is not a looser search. It is a more complete reading. The judgment of what actually belongs to the question has not changed at all. What has changed is only how generously Scroll carries back the things it found.
That wider reach has a price, and we would rather not hide it. Scroll hands the reader about 3,700 tokens at a time, which is close to three times what Tablet returns. More to carry means more to read, so each call does cost more than a lean one would. But the trade is a deliberate one, not something that slipped in by accident. There are questions where being complete matters more than being as cheap as possible, questions where a single missing piece would quietly bend the answer wrong without anyone noticing. For those, paying for the fuller read is not a tax you grudgingly accept on top of the model. It is the whole reason you would reach for Scroll in the first place.
The gain turns up in exactly the place you would expect to find it. On single-session recall, which already sits close to the ceiling, Scroll and Tablet end up within half a point of each other, and that is because when a score is already near perfect there is almost no room left for the fuller read to add anything. It is on the harder ground that the difference truly opens up. Preference inference climbs from 86.0 to 97.3, gathering memory scattered across many sessions rises from 73.8 to 83.9, and reasoning about time moves from 80.8 to 87.1. These are the questions where a single memory left behind is enough to change the answer, and the fuller read is what goes back and brings that missing memory home. Beneath these numbers there is a quieter thing worth stating plainly. The engine had already found the evidence close to 99.6 percent of the time, and it found it for both models in equal measure. So the finding was never what set the two apart. What set them apart was how much of what had been found actually made it all the way to the model. That remaining distance is what we mean when we say delivery, and closing it is the work Scroll was built to do.
We measured all of this in the open, the same way we measured Tablet. We ran the full LongMemEval-S set five separate times, and every one of those runs is laid out just below. A single reader, GPT-5.5, was held constant to answer from the retrieved memories, GPT-4o graded the answers under the official rules, and the exact prompts are published on this same page. The five runs came out at 90.4, 90.0, 91.6, 90.8, and 90.8, for a mean of 90.7% with a standard deviation of 0.5 percent, and all five of them landed at or above 90. We show where Scroll stands firm, and we show just as openly where the road is still long. Gathering memory across many sessions, at 83.9, remains the frontier, and we say so plainly rather than rounding it away. Scroll is the next model to be measured against Tablet 1, and it is held to the same honest measure that the first model set, because only honest numbers earn the kind of trust that lasts.
Verbatim, with nothing paraphrased.
Answer the question using ONLY the retrieved memories below (each
prefixed [date] (rel=relevance)). This question is asked on: {qdate}.
Apply whichever fits:
- 'how long ago' / 'how many days/weeks/months since': compute the
duration relative to the asking date using memory dates.
- If memories give CONFLICTING values for the same fact as of different
dates, give the MOST RECENT value (mention the prior only if asked).
- If the question asks for ADVICE/RECOMMENDATION, first identify this
user's relevant preferences from the memories, then tailor to them.
- For COUNT/TOTAL/LIST questions: ENUMERATE every candidate item in the
memories (including ones mentioned only once or in passing), MERGE
duplicates (same item on different dates = one), then count/sum the
distinct results and state the final number.
- BEFORE the final answer, write an explicit intermediate structure from
the memories: for COUNT/TOTAL/LIST, list every candidate item (incl.
ones mentioned once), merge duplicates (same item on different dates =
one), then count/sum the distinct items; for TEMPORAL /
values-changing-over-time, build a (value, date) timeline. Show the
structure, THEN give the final answer.
If the answer is not in the memories, say you don't know. Answer in the
SAME LANGUAGE as the question.
Memories:
{mems}
Question: {q}
Answer:I will give you a question, the correct answer, and a model's response.
{RULE} Respond with ONLY 'yes' or 'no'.
Question: {q}
Correct answer: {gt}
Model response: {ans}
Is the model response correct?{RULE} is the official LongMemEval rule applied per category:
For temporal-reasoning, an answer is correct if it contains the correct answer or an equivalent, and off-by-one errors in the number of days, weeks, or months are not penalized. For knowledge-update, it is correct if it contains the correct updated answer, and mentioning previous or outdated information is fine as long as the updated answer is present. For single-session-preference, the response need not cover every point in a rubric; it counts as long as it recalls and uses the user's personal information or preference correctly. For all other categories, it is correct if it contains the correct answer, or an equivalent that includes all the intermediate steps to reach it; if it gives only a subset of the required information, it is marked wrong.