Tablet 1 retrieves once and hands back what it found. Tablet 2 keeps that pass and adds a second move: when the first result does not carry the answer, the caller can ask again, and the engine returns memories it has not already delivered rather than the same ones twice, up to three times. No language model runs in that loop. On LongMemEval-S, Tablet 2 scores a mean of 95.7% across three runs, with a standard deviation of 0.4%; the individual runs were 95.2, 96.0, and 96.0. Read by a different frontier model, the same runs score 93.7% (σ 0.3), and both figures are published here for the reason given below. On BEAM 1M, 700 questions over 2.2 million memories, it scores 67.5% across five runs (σ 0.22). The two benchmarks are not on one scale and this page keeps them apart.Tablet 1은 한 번 검색해 찾은 것을 돌려줍니다. Tablet 2는 그 한 번을 그대로 두고 수를 하나 더합니다. 첫 결과에 답이 없으면 호출자가 다시 물을 수 있고, 그러면 엔진은 같은 것을 또 주는 대신 아직 건네지 않은 기억을 최대 세 번까지 돌려줍니다. 이 과정에는 언어 모델이 돌지 않습니다. LongMemEval-S에서 Tablet 2는 3회 측정 평균 95.7%를 기록했고, 표준편차는 0.4%이며, 각 회차 점수는 95.2, 96.0, 96.0이었습니다. 같은 회차를 다른 프런티어 모델이 읽으면 93.7%(σ 0.3)이고, 두 숫자를 모두 싣는 이유는 아래에 적었습니다. 기억 220만 건 위의 700문항인 BEAM 1M에서는 5회 측정 67.5%(σ 0.22)입니다. 두 벤치마크는 같은 자에 놓인 숫자가 아니어서, 이 페이지는 둘을 따로 둡니다.
temperature 0 running BEAM's own judging prompt, the default in the authors' repository. Grading is a separate pass from answering. BEAM fixes neither the reader model nor its prompt, so we name ours: Opus 5, on a five-line prompt.리더는 Claude Opus 5(생각 최대)이고 다섯 회차 내내 고정했습니다. 채점은 gpt-4.1-mini, temperature 0, BEAM 자체 판정 프롬프트 그대로이며 저자 저장소의 기본값입니다. 채점은 답변과 분리된 단계로 돕니다. BEAM 은 리더 모델도 프롬프트도 못박지 않아서 저희 것을 밝힙니다 — Opus 5, 다섯 줄짜리 프롬프트.Every configuration we ran, and the per-category scores behind each one. Nothing is summarised away. The second row is BEAM 1M, a second benchmark on the same engine build: a corpus 40× larger, ten abilities instead of six, and a judge that scores each answer against a checklist rather than as right or wrong. Its numbers are not on one scale with LongMemEval's.돌린 모든 설정과, 각각의 카테고리별 점수입니다. 요약해서 감춘 것은 없습니다. 둘째 줄은 같은 엔진 빌드로 잰 두 번째 벤치마크 BEAM 1M입니다. 코퍼스는 40배 크고, 카테고리 여섯 대신 능력 열 가지로 나뉘며, 평가자가 답을 맞다/틀리다가 아니라 채점 항목표에 대고 잽니다. 그래서 LongMemEval과 같은 자에 놓인 숫자가 아닙니다.
| Category카테고리 | Run 11회 | Run 22회 | Run 33회 | Average평균 |
|---|---|---|---|---|
| Single-session preference단일 세션 선호 | 100.0 | 100.0 | 100.0 | 100.0 |
| Single-session assistant단일 세션 어시스턴트 | 100.0 | 100.0 | 96.4 | 98.8 |
| Temporal reasoning시간 추론 | 97.0 | 97.0 | 97.0 | 97.0 |
| Single-session user단일 세션 사용자 | 95.7 | 94.3 | 95.7 | 95.2 |
| Knowledge update지식 업데이트 | 93.6 | 96.2 | 94.9 | 94.9 |
| Multi-session멀티 세션 | 91.0 | 93.2 | 94.7 | 93.0 |
| Overall전체 | 95.2 | 96.0 | 96.0 | 95.7 |
Mean 95.7% · standard deviation 0.4% · every run published, 0 cherry-picks · 500 / 500 graded in every run. 평균 95.7% · 표준편차 0.4% · 회차 전부 공개, 체리피킹 0 · 매 회차 500 / 500 채점.
Tablet 2 adds one move, and it takes a sentence to state and was hard to build. A memory engine that answers once has to decide, in that one pass, everything it will ever say about a question. It has no way of knowing whether what it found was enough, because nothing tells it afterwards. So it does what any careful system does under that constraint: it returns a little more than needed, in case.Tablet 2가 더한 것은 수 하나입니다. 말로는 한 문장이면 되지만 만드는 일은 어려웠습니다. 한 번에 답하는 기억 엔진은 그 한 번 안에서, 이 질문에 대해 말할 모든 것을 정해야 합니다. 찾아낸 것이 충분했는지 알 길이 없습니다. 뒤에서 알려주는 것이 없기 때문입니다. 그래서 그런 제약 아래의 신중한 시스템이 다 하는 일을 합니다. 혹시 몰라 필요한 것보다 조금 더 돌려줍니다.
Asking again removes that guess. The reader looks at what came back, sees that the answer is not in it, and says so. What returns then is not the same handful shifted by one, but the memories the engine has been keeping track of not having given you yet. It will do that up to three times. Because the engine is the one holding that record, no language model has to sit in the path to make it work, and asking again stays as fast, as steady, and as language-neutral as the first pass.다시 묻는다는 것은 그 짐작을 걷어냅니다. 리더가 돌아온 것을 보고, 답이 그 안에 없음을 확인하고, 그렇다고 말합니다. 그러면 돌아오는 것은 한 칸 밀린 같은 한 줌이 아니라, 엔진이 아직 건네지 않았다고 세어 두고 있던 기억들입니다. 최대 세 번까지 그렇게 합니다. 그 기록을 쥐고 있는 쪽이 엔진이기 때문에, 이것을 돌리자고 경로에 언어 모델을 세울 필요가 없고, 다시 묻는 일도 첫 검색만큼 빠르고 한결같고 언어에 무관합니다.
BEAM 1M · 700 questions · 5 runsBEAM 1M · 700문항 · 5회
Turning verify off is the only change: same engine, same corpus, same reader and judge. Tokens were counted with a tokenizer rather than estimated from character length. It does not pay everywhere — questions whose answer was already in the first pass lose ground, which is why verify is a setting you pass per call.verify 를 끈 것만이 유일한 차이입니다. 엔진도 코퍼스도 리더도 평가자도 같습니다. 토큰은 글자 수에서 어림한 값이 아니라 토크나이저로 셌습니다. 어디서나 이득인 것은 아닙니다. 첫 검색에 이미 답이 있던 문항은 오히려 점수를 잃고, 그래서 verify 는 호출마다 넘기는 설정입니다.
The numbers land where that reasoning predicts, and the comparison that shows it is the one where only a single setting moves. Hold the engine, the corpus, the reader and the judge still, and switch verify off: LongMemEval-S falls from 95.0% to 93.8%, and BEAM 1M from 67.5% to 58.6%. The gain concentrates where one pass is least likely to have gathered everything. Summarization on BEAM 1M needs coverage across sessions rather than one good hit, and it moves 31.4% to 52.5%. It is not free everywhere. Single-session-user questions on LongMemEval-S lose 4.2 points, because asking again when the answer was already in hand adds material that competes with it. Asking again is worth the most exactly where asking once was not enough, and worth less than nothing where it was.숫자는 그 논리가 가리키는 자리에 떨어집니다. 그리고 그것을 보여주는 비교는 설정 하나만 움직이는 쪽입니다. 엔진과 코퍼스와 리더와 평가자를 그대로 두고 verify 만 끄면, LongMemEval-S는 95.0%에서 93.8%로, BEAM 1M은 67.5%에서 58.6%로 내려갑니다. 상승은 한 번의 검색으로 다 모으기 가장 어려운 자리에 몰립니다. BEAM 1M의 요약은 한 번 잘 맞히는 것이 아니라 세션을 두루 훑어야 하는 유형이고, 31.4%에서 52.5%로 움직입니다. 어디서나 공짜인 것은 아닙니다. LongMemEval-S의 단일 세션 사용자 유형은 4.2점을 잃습니다. 답이 이미 손에 있는데 다시 물으면, 그것과 경쟁하는 재료가 붙기 때문입니다. 다시 묻는 일은 한 번으로 부족했던 바로 그 자리에서 가장 값어치가 크고, 충분했던 자리에서는 오히려 손해입니다.
LongMemEval-S asks whether a fact can be found inside a long conversation. BEAM 1M asks something else: whether memory holds its shape at a scale no context window reaches, 2.2 million memories drawn from 74,630 turns. Both are on this page, run by run, and it is the second one that shows the shape of Tablet 2 most plainly. Answering once, it scores 58.6%. Allowed to ask again, 67.5%.LongMemEval-S가 긴 대화 안에서 사실 하나를 찾아낼 수 있는지를 묻는다면, BEAM 1M은 다른 것을 묻습니다. 어떤 컨텍스트 창도 닿지 못하는 규모에서 기억이 제 모양을 지키는가입니다. 74,630턴에서 나온 기억 220만 건입니다. 둘 다 이 페이지에 회차 단위로 실려 있고, Tablet 2의 모양이 가장 또렷하게 보이는 쪽은 두 번째입니다. 한 번에 답하면 58.6%, 다시 물을 수 있으면 67.5%입니다.
Against our own earlier engines the honest comparison is narrower than the headline. Tablet and Scroll are separate tiers rather than successive generations, and the only one where we control both sides is Scroll 1.2 read by the same model: 92.3 to 93.7, a gain of 1.4 points. Those two figures come from separate campaigns rather than paired runs and their intervals overlap, so we report 1.4 as what we measured with the reader held fixed, not as an established difference. Against Tablet 1 there is no controlled comparison at all. Its 85.2% was read by a different model, and the per-question comparison we did produce was withdrawn: a generation claim would need Tablet 1 re-run on the same corpus, the same reader and the same verify budget, and we have not done that.우리 이전 엔진들과의 비교는 헤드라인보다 좁게 잡는 것이 정직합니다. Tablet 과 Scroll 은 이어지는 세대가 아니라 서로 다른 단계이고, 양쪽을 다 통제할 수 있는 비교는 같은 모델이 읽은 Scroll 1.2 하나뿐입니다. 92.3에서 93.7, 1.4점 상승입니다. 다만 두 수치는 짝지어 돌린 것이 아니라 별개 측정에서 나왔고 구간이 겹칩니다. 그래서 1.4점은 리더를 고정하고 잰 값일 뿐 확립된 차이가 아니라고 적습니다. Tablet 1 과는 통제된 비교가 아예 없습니다. 85.2%는 다른 모델이 읽은 값이고, 저희가 만들었던 문항별 비교는 철회했습니다. 세대 간 주장을 하려면 Tablet 1 을 같은 코퍼스, 같은 리더, 같은 verify 예산으로 다시 돌려야 하는데 아직 하지 않았습니다.
Tablet 2 also remembers images. A memory can be an image, and the engine indexes the image itself rather than words about it, so a sentence typed in any language finds it even when the record carries no caption at all. That is the condition worth naming, because it is the one where keyword search has nothing to match: on caption-less images Tablet 2 reaches 91.4% recall@5 and BM25 scores zero, having no text to score. The full protocol is in the paper.Tablet 2는 이미지도 기억합니다. 기억은 이미지일 수 있고, 엔진은 이미지에 대한 말이 아니라 이미지 자체를 색인합니다. 그래서 기록에 캡션이 아예 없어도 어떤 언어로 친 문장이든 그 이미지을 찾습니다. 이 조건을 짚어두는 이유는, 여기가 키워드 검색으로는 대조할 대상 자체가 없는 자리이기 때문입니다. 캡션 없는 이미지에서 Tablet 2는 recall@5 91.4%이고 BM25는 채점할 텍스트가 없어 0입니다. 전체 측정 방법은 논문에 있습니다.
Crossmodal-3600 · 300 images · recall@5Crossmodal-3600 · 이미지 300장 · recall@5
The engine indexes the image rather than words about it, so a sentence in any language finds it with no caption stored. Eleven of the fourteen languages sit above 95, and the two at the bottom are our worst numbers anywhere. BM25 is 0.0 in every one of these cells, not because it scores badly but because a record with no text gives it nothing to score.엔진은 이미지에 대한 말이 아니라 이미지 자체를 색인하므로, 캡션 없이 저장해도 어떤 언어로 친 문장이든 그 이미지를 찾습니다. 14개 언어 가운데 열하나가 95 위에 있고, 바닥의 둘은 우리가 어디서 낸 것보다도 낮은 숫자입니다. BM25 는 이 칸 전부에서 0.0 입니다. 점수가 낮아서가 아니라, 글이 없는 기록에는 점수를 매길 대상 자체가 없기 때문입니다.
An average would hide what matters about that. A retriever can look strong across fourteen languages while being a strong English retriever with thirteen weak ones attached, so the number to read is not the mean but the spread. Tablet 2 varies by 14.0 points across those languages; the two open baselines we could run vary by 27.5 and 27.7, and one of them scores 91.0% on English and 4.7% on Russian from identical image vectors. Avoiding keywords is not the same thing as being language-neutral, and the only way to tell the two apart is to look at the worst language rather than the average one.거기서 정작 중요한 것을 평균이 가립니다. 어떤 검색기는 14개 언어에서 고루 강해 보여도, 사실은 강한 영어 검색기에 약한 열세 언어가 붙어 있는 것일 수 있습니다. 그래서 읽어야 할 숫자는 평균이 아니라 퍼짐입니다. Tablet 2는 그 언어들 사이에서 14.0점 폭으로 흔들리고, 저희가 돌려볼 수 있었던 공개 기준선 둘은 27.5와 27.7입니다. 그중 하나는 동일한 이미지 벡터로 영어 91.0%, 러시아어 4.7%를 냅니다. 키워드를 안 쓴다는 것과 언어를 안 탄다는 것은 같은 말이 아니며, 둘을 가르는 방법은 평균이 아니라 가장 약한 언어를 보는 것뿐입니다.
Two results here work against us, and they belong on the same page as the rest. Retrieval degrades sharply in low-resource languages: Swahili at 53.0% and Telugu at 64.0% are the lowest numbers this model produces anywhere, though the strongest open baseline we could run reaches 6.0% and 5.0% on those same two. And attaching an English caption to an image makes it harder to find in every other language, by 11.4 points on average. That is a defect in something you can buy rather than an artifact of the test, because real stores hold exactly the mixture that produces it. If your users search in more than one language, store images without captions.여기서 우리에게 불리한 결과가 둘 있고, 나머지와 같은 자리에 적어야 할 것들입니다. 저자원 언어에서 검색이 가파르게 나빠집니다. 스와힐리어 53.0%와 텔루구어 64.0%는 이 모델이 어디서 낸 것보다도 낮습니다. 다만 저희가 돌릴 수 있었던 가장 강한 공개 기준선은 같은 두 언어에서 6.0%와 5.0%입니다. 그리고 이미지에 영어 캡션을 붙이면 다른 모든 언어에서 찾기가 어려워집니다. 평균 11.4점입니다. 이는 시험의 인공물이 아니라 살 수 있는 물건의 결함입니다. 실제 저장소가 바로 그 혼합을 담고 있기 때문입니다. 사용자가 여러 언어로 검색한다면 이미지은 캡션 없이 저장하십시오.
The third change is smaller and is aimed at the model rather than at you. An assistant leaning on a store it did not write has no way of knowing how far that memory has been edited underneath it, and the honest thing is to let it ask. Won is where that kind of question lives, and revisions is the first one: how many memories in this store were changed after they were written, against how many were not.세 번째 변화는 더 작고, 사용자보다는 모델을 향합니다. 자기가 쓰지 않은 저장소에 기대는 어시스턴트는 그 기억이 밑에서 얼마나 고쳐졌는지 알 길이 없고, 정직한 방법은 물어볼 수 있게 해 주는 것입니다. Won 은 그런 종류의 질문이 사는 자리이고, revisions 가 그 첫 번째입니다. 이 저장소의 기억 가운데 적힌 뒤에 바뀐 것이 몇 건이고 안 바뀐 것이 몇 건인가입니다.
measured 2026-08-14 · median per call2026-08-14 실측 · 호출당 중앙값
An assistant leaning on a store it did not write should be able to ask how far that memory has been edited underneath it. Won is where those questions live and revisions is the first: two counts, no embedding, no reranker, no language model. It carries no charge, because putting a price on it would make the honest answer the expensive one and callers would stop asking.자기가 쓰지 않은 저장소에 기대는 어시스턴트라면, 그 기억이 밑에서 얼마나 고쳐졌는지 물어볼 수 있어야 합니다. Won 은 그런 질문이 사는 자리이고 revisions 가 그 첫 번째입니다. 개수 둘을 세는 일이라 임베딩도 리랭커도 언어 모델도 돌지 않습니다. 요금은 붙이지 않습니다. 값을 매기면 정직한 답이 비싼 답이 되고, 그러면 아무도 묻지 않게 됩니다.
It answers with counts and nothing else, which is the part that makes it usable. The response is the same size for a store of a hundred memories and a store of a hundred million, so a model can ask it in the middle of a conversation without the answer growing into the context it was trying to protect. It is free for the same reason: putting a price on it would make the honest answer the expensive one, and callers would stop asking. A store where three facts in ten have been replaced deserves less confidence than one nobody has edited, and that is a thing worth knowing before leaning on it.답은 개수뿐이고, 그것이 이 호출을 쓸 만하게 만드는 대목입니다. 기억이 백 건인 저장소든 1억 건인 저장소든 응답의 크기가 같습니다. 그래서 모델이 대화 도중에 물어봐도, 답이 정작 지키려던 컨텍스트를 잡아먹으며 불어나지 않습니다. 무료인 이유도 같습니다. 값을 매기면 정직한 답이 비싼 답이 되고, 그러면 아무도 묻지 않게 됩니다. 열 가지 사실 가운데 셋이 갈아치워진 저장소는 아무도 손대지 않은 저장소보다 덜 믿어야 하고, 그것은 기대기 전에 알아둘 값어치가 있는 사실입니다.
Who ran this.누가 돌렸는가. An internal run. Every number here is self-reported, and so is every number we place it beside.내부에서 돌린 측정입니다. 여기 숫자는 전부 자가 보고이고, 나란히 놓은 다른 숫자들도 마찬가지입니다.
Corpus defects.코퍼스의 흠. 211 of the 74,630 turns were truncated by our ingest harness and one failed to store. Our fault, not the benchmark's, and both stayed in the measured corpus.74,630턴 가운데 211턴이 저희 적재 하네스에서 잘렸고 1턴은 저장에 실패했습니다. 벤치마크가 아니라 저희 잘못이며, 둘 다 측정에 쓴 코퍼스에 그대로 남아 있습니다.