remanence
Try troth
Research & notes

Reading whole papers

Document recall over the full text of research papers from QASPER, a small graded slice.

2026.07.31resulttroth
15of 20 questions answered correctly from papers read in full.

Remembering a conversation and reading a document are different jobs. When you hand troth a paper, a manual or a repository, it is split into overlapping passages, embedded, stored under that document's own scope, and later queried only within it. This note measures that second road.

We took 20 questions from the QASPER development set, one per paper, each paper between eight and thirty-five thousand characters of real full text: title, abstract and every section, not a summary. Fifteen questions ask for spans of the paper, five for a short free-form answer. Unanswerable and yes-or-no questions were left out because a single composed answer grades them badly.

Each paper went through the real ingest path and each answer came back through the real query path, eight passages at a time.

What this does not tell you

  • Twenty items is a small sample. The 95% interval is roughly ±20 points. Read it as a sign the pipeline works end to end, not as a score to compare with anyone else's.
  • The slice was drawn with a fixed seed from filtered papers; the exact filter is written into the result file so it can be drawn again.

A full QASPER run is the next step, and it will be published here the same way, caveats first.

Next: Why we are called RemanenceRead next