One memory, two readers
Conversational recall on a 100-question LongMemEval-S slice, with the whole stack on one machine.
LongMemEval asks an assistant about things said in long, messy conversation histories: a preference mentioned once, a fact that changed twice, how many times something happened across dozens of sessions. It is a fair test of the one thing troth is for, remembering what you settled.
We ran a stratified slice of the first 100 questions of LongMemEval-S, sixteen or seventeen of each type, through the same code path the product uses. For every question a fresh process writes the whole history into memory, digests it without seeing the question, and then answers from what recall brings back. The correct answer is never visible when the answer is written.
Two readers, one memory
The answers were written twice over the same retrieved memory: once by an open-weights 27B model on local hardware, once by Claude Sonnet. Digestion, retrieval and grading stayed local in both cases.
| Question type | n | Local reader | Claude Sonnet |
|---|---|---|---|
| Knowledge update | 16 | 16 | 14 |
| Multi-session | 17 | 12 | 11 |
| Single-session, assistant | 16 | 14 | 14 |
| Single-session, preference | 17 | 12 | 14 |
| Single-session, user | 17 | 14 | 16 |
| Temporal reasoning | 17 | 15 | 15 |
| Total | 100 | 83 | 84 |
The totals sit within a point of each other, and both readers miss the same hard core: counting how often something happened when the history itself describes one event in conflicting ways. That agreement is the result we care about. The memory, not the model, sets the score.
What this does not tell you
- It is a 100-question slice, not the full 500. At this size the noise is roughly ±7 points, so treat smaller differences, including the one between the two readers, as noise.
- The judge is a local open-weights model at temperature 0 using the official per-type prompts, not the GPT-4o used in the paper. A different judge can move absolute numbers.
- The reader sees only the retrieved statements, none of the context a live conversation carries. That isolates the memory and costs some points a real session would recover.
Every number regenerates from our harness against the public dataset, and every per-question verdict is recorded with it.