Measuring memory across languages and images 언어와 이미지를 가로질러 기억을 재다
A memory system has one job: given a situation, return the things worth remembering about it. The usual way to score one is to hand its output to a language model and grade the model's answer, which measures the pair rather than the memory, because a strong reader absorbs a surprising amount of retrieval error. This is the paper without the PDF: what we measured, what moves it, and what it is worth beside anyone else's number. 기억 시스템이 하는 일은 하나입니다. 어떤 상황이 주어지면 그에 대해 기억할 만한 것을 돌려주는 일입니다. 이것을 채점하는 흔한 방법은 그 출력을 언어 모델에 넣고 모델의 답을 채점하는 것인데, 그러면 기억이 아니라 둘의 조합을 재게 됩니다. 강한 리더가 검색 오차를 놀랄 만큼 많이 흡수하기 때문입니다. 이 글은 PDF 를 열지 않은 논문입니다. 무엇을 쟀고, 무엇이 그것을 흔들고, 남의 숫자 옆에 놓았을 때 얼마짜리인지를 다룹니다.
Tablet 2 scores 95.7% on LongMemEval-S, 500 questions, and 67.5% on BEAM-1M, 700 questions over 74,630 conversational turns stored as 2.21 million memories. The engine takes a median of 393 ms per search, and no language model of its own runs anywhere in it. Both figures are quoted with a question-sampling interval rather than a run-to-run spread: repeating either benchmark moves the score by a few tenths, and printing that as ± would claim a precision the question set cannot support. Tablet 2는 500문항의 LongMemEval-S에서 95.7%, 대화 74,630턴을 221만 건의 기억으로 저장한 코퍼스 위의 700문항 BEAM-1M에서 67.5%를 기록했습니다. 검색 한 번에 엔진이 쓰는 시간은 중앙값 393ms이고, 검색 경로 어디에도 자체 언어 모델이 없습니다. 두 수치 모두 회차 간 편차가 아니라 문항 표집 구간과 함께 적습니다. 어느 쪽이든 다시 돌리면 점수가 몇 십분의 일 점 움직이는데, 그것을 ±로 적으면 문항 집합이 뒷받침하지 못하는 정밀도를 주장하는 것이 됩니다.
What actually moves the number?그 숫자는 무엇에 따라 움직이나
Hold the engine, the corpus, the retrieval settings and the judge completely still, and change one thing. Changing only how many times the caller may ask again moves BEAM-1M by 8.9 points. Changing only which model reads the retrieved memories moves LongMemEval-S by 2.0. Neither is stated in the reports we would be compared against. 엔진과 코퍼스와 검색 설정과 판정자를 완전히 고정해두고 한 가지만 바꿔봅니다. 호출자가 몇 번까지 다시 물을 수 있는지만 바꾸면 BEAM-1M이 8.9점 움직입니다. 검색된 기억을 어느 모델이 읽는지만 바꾸면 LongMemEval-S가 2.0점 움직입니다. 우리가 비교 대상이 될 리포트들은 두 설정 어느 쪽도 밝히지 않습니다.
One setting against the distance the comparison is about설정 하나와, 그 비교가 다루는 거리
The second bar is not a setting. It is the entire distance between the highest published system and ours, and one knob that nobody states is worth more than it. That is why the chart below is a placement and not a ranking.두 번째 막대는 설정이 아닙니다. 공개된 가장 높은 시스템과 우리 사이의 거리 전체이고, 아무도 밝히지 않는 손잡이 하나가 그보다 큽니다. 아래 그림을 순위가 아니라 배치라고 부르는 이유입니다.
Where does it sit beside everyone else?남들과 나란히 놓으면 어디쯤인가
Third of seven, on the numbers their authors publish. We show the comparison because leaving it out would be its own kind of claim, and we call it a placement because the chart above says what a ranking would be worth: the gap between first and us is 7.5 points, and a single unstated setting is worth 8.9. 저자들이 공개한 숫자로 보면 일곱 중 셋째입니다. 비교를 싣는 이유는, 빼는 것 자체가 또 하나의 주장이 되기 때문입니다. 그리고 이것을 순위가 아니라 배치라고 부르는 이유는 위의 그림이 말해 줍니다. 1위와 우리 사이가 7.5점인데, 아무도 밝히지 않는 설정 하나가 8.9점짜리입니다.
Retrieved August 2026 · every number including ours is self-reported2026년 8월 기준 · 우리 것을 포함해 전부 자가 보고
No two rows are known to share a reader, and the re-ask budget is stated in none of them. Since one of those settings is worth 8.9 points on its own, read the ordering as where the systems sit rather than which is better. Of the systems above, only Mem0 states a token count: about 6,900 per query against our 2,484.어느 두 줄이 같은 리더를 썼는지 알 수 없고, 되묻기 예산은 어느 줄에도 적혀 있지 않습니다. 그 설정 하나가 8.9점짜리이므로, 이 순서는 어느 쪽이 낫다가 아니라 각자 어디쯤 있다로 읽어야 합니다. 위에서 토큰 수를 밝힌 것은 Mem0 뿐이고, 질의당 약 6,900 토큰으로 우리 2,484 와 견줍니다.
The comparison we can actually control is against our own engines, and it is smaller than the headline. Read by the same model, Scroll 1.2 scores 92.3 and Tablet 2 scores 93.7, a gain of 1.4 points. Those come from separate campaigns rather than paired runs and their intervals overlap, so we report 1.4 as what we measured with the reader held fixed, not as an established difference. Against Tablet 1 there is no controlled comparison at all: its 85.2% was read by a different model, and the per-question comparison we produced was withdrawn. 우리가 실제로 통제할 수 있는 비교는 우리 엔진끼리이고, 헤드라인보다 작습니다. 같은 모델이 읽으면 Scroll 1.2가 92.3, Tablet 2가 93.7로 1.4점 차입니다. 다만 짝지어 돌린 것이 아니라 별개 측정에서 나왔고 구간이 겹치므로, 1.4점은 리더를 고정하고 잰 값일 뿐 확립된 차이가 아니라고 적습니다. Tablet 1과는 통제된 비교가 아예 없습니다. 85.2%는 다른 모델이 읽은 값이고, 저희가 만들었던 문항별 비교는 철회했습니다.
Can you see the retrieval on its own?검색만 따로 볼 수 있나
Almost. One configuration comes close enough to be worth pointing at. Switch verify off entirely and drop the reader to medium effort, so the engine gets exactly one retrieval and the model reading it is the weaker of the two we tried. It scores 93.8, which edges the stronger reader running three retrieval passes, and it hands that reader 30% fewer memories to work with. That is the closest thing in the paper to a measurement of retrieval quality with the reader held still. 거의 됩니다. 한 설정이 가리킬 만큼은 가깝습니다. verify 를 완전히 끄고 리더를 medium 노력으로 낮추면, 엔진은 검색을 정확히 한 번만 하고 그것을 읽는 모델은 우리가 시도한 둘 중 약한 쪽이 됩니다. 그 설정이 93.8을 냅니다. 검색을 세 번 도는 더 강한 리더를 앞서면서, 그 리더에게 건네는 기억은 30% 적습니다. 논문에서 리더를 세워둔 채 검색 품질을 잰 것에 가장 가까운 숫자입니다.
What does one number across every language buy?모든 언어에서 같은 숫자라는 게 무엇을 벌어주나
The claim is only worth something measured against the thing it rules out. So we built a corpus where the answer is checkable by eye — thirty images, five colours by six shapes, so that “red circle” has exactly one right answer — and crossed seven store conditions with ten query languages to get seventy cells. 그 주장은 그것이 배제하는 대상과 나란히 재었을 때만 값이 있습니다. 그래서 눈으로 확인할 수 있는 코퍼스를 만들었습니다. 색 다섯에 모양 여섯으로 이미지 서른 장을 만들어 "빨간 원"의 정답이 정확히 하나가 되게 하고, 저장 조건 일곱 가지와 질의 언어 열 가지를 교차해 일흔 칸을 얻었습니다.
Controlled corpus · 70 store-language × query-language cells · 30 queries each통제 코퍼스 · 저장 언어 × 질의 언어 70칸 · 칸마다 질의 30개
BM25 was configured as strongly as we could make it: standard parameters, two tokenizers run over every cell, the better of the two reported per cell. Its median across the seventy is 0.0. The only off-diagonal cells it wins are between relatives — Japanese to Chinese, which share characters, and English to German and Spanish, where cognates give a bigram tokenizer partial overlap.BM25 는 저희가 만들 수 있는 한 강하게 구성했습니다. 표준 파라미터에, 토크나이저 둘을 모든 칸에 돌리고 칸마다 더 나은 쪽을 적었습니다. 70칸 중앙값은 0.0 입니다. 대각선 밖에서 점수가 나는 칸은 친척뿐입니다. 한자를 공유하는 일본어와 중국어, 그리고 동계어 덕에 바이그램 토크나이저가 일부 겹치는 영어에서 독일어·스페인어입니다.
The store conditions mattered more than they looked, and getting them wrong once produced a confident, wrong answer. Our first version of this experiment mixed captioned and caption-less images in one store; the captioned ones crowded the five result slots and pushed the others out, and we concluded that text queries cannot find caption-less images at all. Zero percent. Separating the stores reversed it. We report the mixed condition as a result rather than hiding it, because real customer stores are mixed. 저장 조건은 보기보다 중요했고, 한 번 잘못 잡아 확신에 찬 오답을 냈습니다. 이 실험의 첫 판은 캡션 있는 이미지와 없는 이미지를 한 저장소에 섞었습니다. 캡션 있는 쪽이 결과 다섯 자리를 차지해 나머지를 밀어냈고, 우리는 텍스트 질의로는 캡션 없는 이미지를 아예 못 찾는다고 결론지었습니다. 0퍼센트라고요. 저장소를 분리하니 뒤집혔습니다. 섞인 조건도 감추지 않고 결과로 싣습니다. 실제 고객 저장소가 섞여 있기 때문입니다.
Does a dense system come out language-neutral?밀집 방식이면 언어를 안 타게 되나
No, and this is the control we thought mattered most, because no serious competitor retrieves images with keywords anyway. We ran two open baselines that also avoid term matching, built so that only the text side differs between them: the same image encoder, one with the original English-trained text tower and one distilled to cover more than fifty languages. Across those fourteen languages Tablet 2 averages 91.4% recall@5; what separates it from the baselines is not that average but how little it moves between languages. 아닙니다. 그리고 이것이 가장 중요하다고 본 대조군입니다. 어차피 이미지를 키워드로 검색하는 진지한 경쟁자는 없기 때문입니다. 마찬가지로 용어 대조를 피하는 공개 기준선 둘을 돌렸고, 둘 사이에 텍스트 쪽만 달라지도록 짰습니다. 이미지 인코더는 같고, 하나는 원래의 영어 학습 텍스트 타워, 하나는 오십 개 넘는 언어를 덮도록 증류한 것입니다. 그 14개 언어에서 Tablet 2 는 recall@5 평균 91.4%입니다. 다만 기준선들과 갈리는 것은 그 평균이 아니라 언어 사이에서 얼마나 덜 흔들리는가입니다.
Crossmodal-3600 · 300 images, no captions · 14 languagesCrossmodal-3600 · 캡션 없는 이미지 300장 · 14개 언어
Both rows retrieve from the identical stored image, so the second chart is not a vision failure: it is a text encoder that was never asked to cover Russian. Swapping in a multilingual encoder lifts the mean from 19.4% to 68.5% and is level across nine of the fourteen, then falls to 5.0% on Telugu and 6.0% on Swahili. Flat over the languages you trained on is not the same as flat.두 줄 모두 동일한 이미지 벡터에서 검색하므로, 두 번째 그림은 시각 쪽 실패가 아닙니다. 러시아어를 다루라고 요구받은 적이 없는 텍스트 인코더입니다. 다국어 인코더로 바꾸면 평균이 19.4%에서 68.5%로 오르고 14개 중 아홉에서 고르지만, 텔루구어 5.0%와 스와힐리어 6.0%로 떨어집니다. 학습한 언어들 위에서 평평한 것과 평평한 것은 다릅니다.
Does delivering more keep working?더 많이 주면 계속 좋아지나
Up to a point, and then it stops paying. Turning on context expansion, which attaches neighbouring material to each retrieved memory, scores highest of the three configurations we ran. We do not present it as our result, and the chart says why. 어느 지점까지는 그렇고, 그다음부터는 값을 못 합니다. 검색된 기억마다 주변 자료를 붙이는 맥락 확장을 켜면, 우리가 돌린 세 설정 가운데 가장 높은 점수가 나옵니다. 그것을 우리 결과로 내세우지 않는데, 이유는 그림이 말해 줍니다.
BEAM-1M · same engine, same corpus, same reader and judgeBEAM-1M · 엔진·코퍼스·리더·평가자 모두 동일
The first 8.9 points cost 3.2× the context. The last 2.8 cost 6.7× more on top of that, and that run is preliminary: 681 of 700 questions completed. It is also not uniformly better. Three question types score worse with expansion on than without it, and they are the three where the answer depends on telling similar records apart, which is exactly what more similar records makes harder. The open question is not how much context to deliver but to which questions.처음 8.9점은 맥락 3.2배를 썼습니다. 마지막 2.8점은 그 위에 6.7배를 더 씁니다. 그리고 그 회차는 예비 측정입니다. 700문항 중 681문항만 끝났습니다. 고르게 좋아지지도 않습니다. 세 유형은 확장을 켜면 오히려 점수가 내려가는데, 비슷한 기록들을 가려내야 답이 되는 바로 그 세 유형이고, 비슷한 기록을 더 주는 일이 어렵게 만드는 지점입니다. 남은 질문은 맥락을 얼마나 줄까가 아니라 어느 질문에 줄까입니다.
Where is it bad?어디서 못하나
Retrieval degrades sharply in low-resource languages: 53.0% recall@5 for Swahili and 64.0% for Telugu, the lowest numbers we produce anywhere, though the strongest open baseline we could run reaches 6.0% and 5.0% on those same two. Attaching an English caption to an image makes it harder to find in other languages, by 11.4 points on average across the fourteen. That is a defect in something people can buy rather than an artifact of the test, because real stores hold exactly that mixture. 저자원 언어에서 검색이 가파르게 나빠집니다. 스와힐리어 recall@5 53.0%, 텔루구어 64.0%로 우리가 어디서 낸 것보다도 낮습니다. 다만 우리가 돌릴 수 있었던 가장 강한 공개 기준선은 같은 두 언어에서 6.0%와 5.0%입니다. 이미지에 영어 캡션을 붙이면 다른 언어에서 찾기가 어려워집니다. 14개 언어 평균 11.4점입니다. 이는 시험의 인공물이 아니라 사람들이 살 수 있는 물건의 결함입니다. 실제 저장소가 바로 그 혼합을 담고 있기 때문입니다.
Our worst BEAM-1M type is event ordering at 23.6%, and we think it measures something other than ordering. The rubric expects ten topics; our answers give ten individual events. Both are in chronological order and both have ten items, but they sit at different levels of granularity, so the matcher fails to pair them and the correlation goes negative. We could raise the score by reshaping answers to the rubric's granularity. We did not, because that is answering to the answer key rather than to the question. BEAM-1M 에서 우리 최악 유형은 사건 순서로 23.6%인데, 이것이 재는 것은 순서 능력이 아니라고 봅니다. 채점표는 열 개의 주제를 기대하고 우리 답은 열 개의 개별 사건을 냅니다. 둘 다 시간순이고 둘 다 열 항목이지만 낟알의 크기가 달라서, 짝짓기가 실패하고 상관계수가 음수로 갑니다. 답을 채점표의 낟알에 맞춰 다시 쓰면 점수를 올릴 수 있습니다. 하지 않았습니다. 그것은 질문이 아니라 정답지에 답하는 일이기 때문입니다.
Whose limit is a weak language?약한 언어는 누구의 한계인가
From the outside these look identical: a language scores badly, and the component we did not write is the obvious suspect. Twice we had to decide which it was, and we got it wrong once. One setting omitted on the way into a stage of our own retrieval cost 37 points of Korean accuracy on short queries while leaving nine other languages untouched. We had written down the opposite conclusion, that the limit was inherent to that stage, and held it for two weeks. 밖에서 보면 둘은 똑같이 생겼습니다. 어떤 언어의 점수가 나쁘고, 우리가 만들지 않은 부품이 자연스러운 용의자가 됩니다. 그것을 가려야 하는 일이 두 번 있었고, 한 번은 틀렸습니다. 우리 검색의 한 단계로 들어가는 길에 빠뜨린 설정 하나가 짧은 질의에서 한국어 정확도를 37점 깎았고, 다른 아홉 언어는 건드리지 않았습니다. 우리는 그 반대의 결론을, 그 한계가 그 단계에 내재한 것이라고 적어두고 2주 동안 유지했습니다.
The low-resource case really was the component's. Separating the two took the same procedure both times: measure the stage on its own, with everything around it removed, and see which number you reproduce. A limit you have misattributed is worse than one you have not found, because you stop looking for it. 저자원 언어 쪽은 정말로 부품의 한계였습니다. 두 번 다 같은 절차로 갈렸습니다. 그 단계만 따로, 주변을 전부 걷어내고 재어서 어느 숫자가 재현되는지 봅니다. 잘못 귀속한 한계는 못 찾은 한계보다 나쁩니다. 찾기를 멈추게 되기 때문입니다.
Who wrote this누가 썼는가
This is our own product and the paper was written by the person who builds it. Every number in it is self-reported, and so is every number we place it beside. What we can offer instead of independence is the procedure: the settings that move the score are published, the benchmarks are public datasets, every run is in the paper, and the work can be re-run rather than believed. 이것은 우리 제품이고, 논문은 그것을 만든 사람이 썼습니다. 논문의 모든 숫자가 자가 보고이며, 나란히 놓은 다른 숫자들도 마찬가지입니다. 독립성 대신 내놓을 수 있는 것은 절차입니다. 점수를 움직이는 설정을 공개했고, 벤치마크는 공개 데이터셋이며, 모든 회차가 논문에 있고, 이 작업은 믿는 것이 아니라 다시 돌려볼 수 있습니다.
Footnotes각주
Intervals, not error bars.오차 막대가 아니라 구간입니다. 95.7% carries [93.4, 97.1] and 67.5% carries [64.8, 70.2]. Those are question-sampling intervals: how much the score could move if you drew a different set of questions of the same size. Repeating a benchmark on the same questions moves it far less, and quoting that smaller number as ± would look more precise while saying less.95.7%는 [93.4, 97.1], 67.5%는 [64.8, 70.2]를 답니다. 문항 표집 구간입니다. 같은 크기의 다른 문항 집합을 뽑았을 때 점수가 얼마나 움직일 수 있는가입니다. 같은 문항으로 다시 돌릴 때의 변동은 그보다 훨씬 작고, 그 작은 값을 ±로 적으면 더 정밀해 보이면서 말하는 바는 적어집니다.
The expansion run is preliminary.확장 회차는 예비 측정입니다. 70.3% is over the 681 questions that completed, against 700 for the headline configuration. It is reported as headroom that exists, not as a result.70.3%는 끝난 681문항 기준이고, 헤드라인 설정은 700문항입니다. 결과가 아니라 여지가 있다는 표시로 싣습니다.
The two benchmarks use different judges.두 벤치마크의 판정자가 다릅니다. Each ships its own default and we changed neither, so 95.7 and 67.5 are not comparable to each other. They measure different corpora under different scoring rules.각 벤치마크가 자기 기본값을 싣고 있고 저희는 어느 쪽도 바꾸지 않았습니다. 그래서 95.7 과 67.5 는 서로 비교할 수 없습니다. 다른 코퍼스를 다른 채점 규칙으로 잰 값입니다.
Grading BEAM-1M is expensive.BEAM-1M 채점은 비쌉니다. Nugget scoring calls the judge once per rubric item, about 3.45 times per question. Five runs took 38,662 judge calls and 25.3M tokens, more than the question count would suggest and worth knowing before anyone plans a reproduction.항목 채점은 채점표 항목마다 판정자를 한 번씩 부르며, 문항당 평균 3.45회입니다. 5회 측정에 판정 호출 38,662회, 25.3M 토큰이 들었습니다. 문항 수만 보고 짐작하는 것보다 크고, 재현을 계획하는 사람이 미리 알아둘 값어치가 있습니다.
Read the full paper on arXiv논문 전문 읽기 (arXiv) →
Was this report useful?이 리포트가 도움이 되었나요?
We publish every run and our own conflict of interest so this work can be checked. If something here is wrong, we would rather hear it. 이 작업이 검증받을 수 있도록 모든 실행과 이해충돌을 공개합니다. 여기 잘못된 것이 있다면, 듣는 편이 낫습니다.