Tablet 1 is the first publicly available WOS model: the memory engine you can call today. We publish its results the same way we intend to publish every model that follows it: the full protocol, every run, nothing cherry-picked. On LongMemEval-S, a long-context memory benchmark, Tablet 1 scores a mean of 85.2% across five runs, with a standard deviation of 1.1%; the individual runs were 86.2, 84.2, 85.0, 83.8, and 86.6.
temperature 0, returning a plain yes / no under the official LongMemEval per-category rules.The raw scores behind the 85.2%.
| Category | Run 1 | Run 2 | Run 3 | Run 4 | Run 5 | Average |
|---|---|---|---|---|---|---|
| Single-session user | 98.6 | 100.0 | 98.6 | 98.6 | 98.6 | 98.9 |
| Single-session assistant | 94.6 | 96.4 | 96.4 | 96.4 | 98.2 | 96.4 |
| Knowledge update | 92.3 | 89.7 | 93.6 | 88.5 | 92.3 | 91.3 |
| Preference inference | 90.0 | 86.7 | 83.3 | 80.0 | 90.0 | 86.0 |
| Temporal reasoning | 82.0 | 78.9 | 81.2 | 78.9 | 82.7 | 80.7 |
| Multi-session | 75.9 | 72.2 | 72.2 | 73.7 | 75.2 | 73.8 |
| Overall | 86.2 | 84.2 | 85.0 | 83.8 | 86.6 | 85.2 |
Mean 85.2% · standard deviation 1.1% · all 5 runs published, 0 cherry-picks.
Curves show the shape of the distribution around the measured median; exact spread varies by workload.
Remembering is a quiet thing. To remember well is not to recall a great deal, but to bring back the one thing that fits, quietly, at the moment it is needed. Human memory works that way. We do not unfold every day we have lived all at once; we surface the single piece this conversation needs and let the rest rest. Most systems do the opposite. As a conversation grows, they gather everything that came before and pour all of it in front of the model, every single time. The more that accumulates, the slower and costlier it becomes, and the focus blurs exactly where it matters most. Tablet 1 began from the other side. Not more, but enough.
Tablet 1 does not look at whether words overlap. It looks at whether meaning meets. Matching by words is bound to the surface of language, so the moment the same idea is written in different words, or in a different language, recall quietly collapses. Searching by meaning removes that wall. Ask in English, ask in Japanese, ask in Chinese, and it remembers with the same depth, and it does not mind when a single person's memory holds several languages at once. To answer with the same weight for anyone, in any language. That is the promise this engine set out to keep from the start. The method we keep within. We will show you what it does; how it does it stays ours.
What comes back is always small, and always the same. Whether you have kept a thousand memories or a million, a single call returns a handful of steady size. No language model sits in the path of retrieval. So the same records always return the same memories. Ask today and ask again tomorrow, and it does not waver. To entrust something is, in the end, to trust this constancy. Fast, light, and possible to anticipate. Past any flourish of words, it is this steadiness that lasts.
Even when the words differ, Tablet 1 follows the grain between one memory and the next and brings the surrounding context along with it. And when it is not sure, rather than offering a memory that is plausible but wrong, it stays quiet. An empty space can be filled again; a wrongly filled one only leads a person further astray. Not pretending to know what it does not. That is where trust begins. It reads time as well. It tells when something happened from when it was written down, returns things in order when asked, and conveys how long ago a thing was without handing the model an exact date to invent from.
All of this, we measured in the open. LongMemEval-S, a long-term memory benchmark, tests six kinds of remembering across 500 questions: facts within a single session, facts the assistant left behind, updating information that has changed, inferring preferences, reasoning about time, and gathering pieces scattered across many sessions. We ran it five times under the same conditions; the answers were read and written by Claude Opus 4.8, and the grading was done by GPT-4o. There is a reason we placed a strong reader at the end. It lets the quality of retrieval show plainly in the quality of the answer. Because the engine returns the same memories for the same records, the small movement between runs comes from the model writing the answer, not from retrieval.
The result was a mean of 85.2%, with a standard deviation of 1.1%. The five runs scored 86.2, 84.2, 85.0, 83.8, and 86.6, and rather than holding up the best of them, we laid all five out as they were. On the easier ground, like recall within a single session, it reaches close to 99%; on the hardest, gathering memory scattered across many sessions, it still holds near 74%. We hid neither where it stands firm nor where the road is still long. Only honest numbers earn trust that lasts.
The largest difference shows in the quietest place. The past conversation a single question carries averages around 140,000 tokens. Of that vast record, what Tablet 1 actually handed to the model was a median of about 1,200 tokens at a time, and never more than 1,700. Less than one percent of what had accumulated. The rest was not thrown away. It was kept whole, while only the piece needed in that moment was passed along. So even as memory swells into the millions, the size of the input going to the model stays nearly where it was. That cost does not balloon as memory grows is a rare virtue in a tool you mean to keep by your side for a long time. Retrieval time on the engine alone held at a median of 320 milliseconds, staying between 170 and 580 even as questions grew complex.
And this is a beginning. Tablet is our first model, and every model that follows will be measured and weighed against this one, Tablet 1. Because the first model was honest, every promise that follows is held to the same measure.
Verbatim, nothing paraphrased.
Answer the question using ONLY the retrieved memories below (each is
prefixed with its [date]). This question is being asked on: {qdate}.
Apply whichever of these fits the question:
- For any 'how long ago' / 'how many days/weeks/months since' question,
compute the duration relative to the asking date above (not any other
today), using the memory dates.
- If the memories give CONFLICTING values for the same fact (different
values as of different dates), mention BOTH and note which is more recent.
- If the question asks for ADVICE or a RECOMMENDATION, first identify this
user's relevant preferences, interests, and past choices from the
memories, then tailor your answer to them (not generic advice).
- Otherwise, answer the factual question concisely and directly.
If the answer is not in the memories, say you don't know. Answer in the
SAME LANGUAGE as the question.
Memories:
{mems}
Question: {q}
Answer:I will give you a question, the correct answer, and a model's response.
{RULE} Respond with ONLY 'yes' or 'no'.
Question: {q}
Correct answer: {gt}
Model response: {ans}
Is the model response correct?{RULE} is the official LongMemEval rule applied per category:
For temporal-reasoning, an answer is correct if it contains the correct answer or an equivalent, and off-by-one errors in the number of days, weeks, or months are not penalized. For knowledge-update, it is correct if it contains the correct updated answer; mentioning previous or outdated info is fine as long as the updated answer is present. For single-session-preference, the response need not cover every point in a rubric; it counts as long as it recalls and uses the user's personal information or preference correctly. For all other categories, it is correct if it contains the correct answer, or an equivalent that includes all the intermediate steps to reach it; if it gives only a subset of the required information, it is marked wrong.