remanence
Try troth
Research & notes

One memory, two readers

Conversational recall on a 100-question LongMemEval-S slice, with the whole stack on one machine.

2026.08.31resulttroth
83of 100 questions, with every part of the stack on one machine.

LongMemEval asks an assistant about things said in long, messy conversation histories: a preference mentioned once, a fact that changed twice, how many times something happened across dozens of sessions. It is a fair test of the one thing troth is for, remembering what you settled.

We ran a stratified slice of the first 100 questions of LongMemEval-S, sixteen or seventeen of each type, through the same code path the product uses. For every question a fresh process writes the whole history into memory, digests it without seeing the question, and then answers from what recall brings back. The correct answer is never visible when the answer is written.

Two readers, one memory

The answers were written twice over the same retrieved memory: once by an open-weights 27B model on local hardware, once by Claude Sonnet. Digestion, retrieval and grading stayed local in both cases.

Question typenLocal readerClaude Sonnet
Knowledge update161614
Multi-session171211
Single-session, assistant161414
Single-session, preference171214
Single-session, user171416
Temporal reasoning171515
Total1008384

The totals sit within a point of each other, and both readers miss the same hard core: counting how often something happened when the history itself describes one event in conflicting ways. That agreement is the result we care about. The memory, not the model, sets the score.

What this does not tell you

  • It is a 100-question slice, not the full 500. At this size the noise is roughly ±7 points, so treat smaller differences, including the one between the two readers, as noise.
  • The judge is a local open-weights model at temperature 0 using the official per-type prompts, not the GPT-4o used in the paper. A different judge can move absolute numbers.
  • The reader sees only the retrieved statements, none of the context a live conversation carries. That isolates the memory and costs some points a real session would recover.

Every number regenerates from our harness against the public dataset, and every per-question verdict is recorded with it.

Next: Reading whole papersRead next