Research연구

We build long-term memory for AI. Our research covers how models remember, recall, and stay continuous over time—on public benchmarks, with every run reported. 우리는 AI를 위한 장기 기억을 만듭니다. 모델이 어떻게 기억하고 회상하며, 시간이 지나도 연속성을 잃지 않는지를 연구합니다. 공개 벤치마크 위에서, 모든 실행을 공개해 측정합니다.

Evaluation평가 Aug 19, 20262026년 8월 19일 Featured주요 연구
Wontopos Tablet 2: measuring multilingual and multimodal memory retrieval Wontopos Tablet 2: 다국어·멀티모달 기억 검색 실측

Tablet 2 scores 95.7% on LongMemEval-S and 67.5% on BEAM-1M, every run published, and it finds a photograph stored with no caption from a sentence in any language, where keyword search has nothing to score at all. Most of the paper is about how little a score means on its own, and it reports three findings that run against us at the same weight as the rest. Tablet 2는 LongMemEval-S에서 95.7%, BEAM-1M에서 67.5%를 기록했고 모든 회차를 공개합니다. 캡션 없이 저장한 사진도 어떤 언어로 친 문장으로 찾아냅니다. 키워드 검색으로는 점수를 매길 대상 자체가 없는 조건입니다. 논문의 대부분은 점수 하나가 그것만으로 얼마나 적은 것을 말하는지에 쓰고, 우리에게 불리한 결과 셋도 같은 무게로 싣습니다.

Read the paper →논문 읽기 →

Open source오픈소스

Every number we publish can be reproduced. The tools we use to measure will be opened here. 우리가 발표하는 모든 수치는 재현할 수 있습니다. 측정에 쓰는 도구는 이곳에서 공개할 예정입니다.

WMB-100K

100,000 turns and 2,708 questions, with the harness and scoring code to re-run it yourself.10만 턴, 2,708문항. 직접 다시 돌려볼 수 있는 하네스와 채점 코드까지 있습니다.

Work in progress. The dataset, the scoring and the harness are still rough, so treat results as preliminary.아직 만드는 중입니다. 데이터셋과 채점과 하네스가 거칠어서, 결과는 잠정치로 봐주세요.

GitHub →
Rust · Python