Long-term memory for AI agents.
WOS is a memory API. You store a user's memories once, then recall only the relevant ones for each query and pass them to your model's prompt.
Recall quality is the same across languages, measured at 95.2% recall@5 over 70 store-language by query-language pairs. Each query returns a small, bounded context no matter how much you have stored, and no model is ever run over your stored memories.
Core operations
store- save a memory for a user.recall- get the relevant memories for a query. This is the main call.search- raw semantic search over stored memories.supersede- update or replace a memory that is out of date.forget- delete a single memory or an entire user (GDPR).
Beyond the context window.
WOS recalls from histories of 1.4M tokens - far larger than any LLM context window - and still hands back a tight ~1,470-token slice.
Your agent's memory isn't capped by what fits in a prompt. It keeps everything and retrieves only what matters, no matter how large the history grows.
Private, and yours.
Your data stays in your store. We never train on it, view it, or reuse it - we only organize it so you can retrieve it.
- BYOK. Your LLM key is sent per request and never stored.
- Isolated. Memories are scoped per workspace, then per store (
user_id). - GDPR delete. One call wipes a user - every memory, image and revision, with nothing kept behind.
Why WOS
- Three models, one lineage. WOS models are named for how people have kept knowledge through history - Tablet, Scroll, Book. Stone, scroll, bound book: each one does more for your agent
- Pay us $2. Save many times that on your LLM. WOS feeds your LLM ~1,000 tokens per query - a bounded, relevant slice - instead of stuffing the full history into every prompt. The gap is enormous, and it
- Every language, the same accuracy. Recall quality is the same whether your users write in 日本語, 中文, Español, or English. We measure it as a grid of 70 store-language by query-language pairs and
- No model runs over your memories. Nothing you send is rewritten, and the engine is cheap, fast and deterministic. A model is never run over your stored memories. Tablet uses no model at all
- 67.5%, measured and reproducible. 67.5% on BEAM 1M, averaged over 5 independent runs (σ 0.22%, none cherry-picked), graded by gpt-4.1-mini with the benchmark's own judging prompt.
- Two token rates per model, plus $0.0001 per request. Per million tokens plus a flat $0.0001 per request, pay as you go. No subscription, no storage rent, no memory caps. You pay when your agent writes or reads
For developers
- Three calls: store, recall, answer. One API. The recall() call returns short-term, long-term, and surrounding context in a single round-trip, ready to drop into your prompt.
- Your first recall in 5 minutes. One key, one install line, three calls - your agent has memory. Every snippet on this page was actually run; responses are shown verbatim.
- Stores - create, list, delete. A store is the user_id you read and write under - one isolated memory space per end-user, agent, or topic. Stores are explicit : create one before you store
- Repeated recall, at a tenth of the price. Opt in per request and WOS caches the search result under its query text, with the same prefix rules as LLM prompt caching. While the cache is warm, a
- Memory that knows who said it. People remember by person: what Bob promised, what you said you would do. Tag each memory with a speaker and your agent does the same, on every Tablet and
- Images A memory can carry an image. The engine indexes the image, so a text query in any language matches it even when the record has no caption, title or alt text.
- verify verify lets a search run additional passes. Each pass excludes what earlier passes returned, so a second pass reaches memories the first did not.
- lineage The chain of edits behind one memory, oldest first. revisions says how much a store moved; this says what happened to one fact.
- Your memory inside every AI tool One command gives Claude Code, Claude Desktop, Cursor, or any MCP host a long-term memory backed by your WOS account. No integration code - the agent gets
- Engrams Callable recall tools your model can invoke - each one a different retrieval strategy over the same memory. Use one, or run several at once.
- Every endpoint, one base URL. No SDK required - any HTTP client works. Base URL https://api.wontopos.com , auth via the X-API-Key header, JSON in and out. Memory ops are POST. Stores use
- Same features for everyone. Tiers only raise your limits. Every tier runs the full engine - same recall quality, same languages, every method. Tiers advance automatically through Tier 5 as your cumulative credit
- When something goes wrong. Errors come back as a JSON envelope with a stable type , a human message, and a request_id you can send us when reporting an issue.
- Won is for whoever reads the memory. Most of this API answers with memories. Won answers about them: how much a store has been revised, and how far it can be trusted. Read-only, free, and
- Python - every method, three groups. Python - every method, three groups. Write, read, delete. Every example below was run against the live API on 2026-08-01; responses are verbatim.
- TypeScript - every method, three groups. TypeScript - every method, three groups. Write, read, delete. Every example below was run against the live API on 2026-08-01; responses are verbatim.
- Rust - every method, three groups. Rust - every method, three groups. Write, read, delete. Every example below was run against the live API on 2026-08-01; responses are verbatim.
- curl - no install, same methods. No SDK to install - any HTTP client works. Set your key once and call the same endpoints the SDKs wrap. Base URL https://api.wontopos.com , auth via
- Time_awareness A delivery form - pick it per call. Pass form - memoir or archive - with any call on a form-capable model (Scroll 1.2 and up), and the response comes back
- deep_recall Multi-hop recall. Follows links between memories, bringing in context a single search would miss. Best when memories reference each other
- timeline Time-ordered recall. Returns memories sorted newest-first by when the event happened , not by relevance. For "when did X", history, and sequence questions
- gather Broad gather. A wider net than deep_recall. Use it to pull in everything related to a person, project, or topic in one call. Returns up to ~18.
- equilibrium Drift correction. As a session runs long, replies can start to loop, flatten, or circle whatever the last stretch was about. This brings the range back. Use
- tone_stabilizer Its own voice. Long sessions pull an assistant off its register: replies stretch, turn into reports, or take on the mood of the last stretch. Ordinary
- What this key has spent, and what is left usage answers "can I keep going?". It returns this key’s own lifetime cost, the workspace it belongs to over a window, what each store cost inside that
- How much of a store has been rewritten revisions answers revised out of total : how many memories in a store were altered after they were written. It is worth asking before leaning on memory for
More
- Add-on The complete machine-readable map of the API - every endpoint, request, response, and error.