How Small a Model Can Read Your Documents
The decision memo version of “Reference, Retrieve, Resolve: How Small a Language Model Can Read a Reference in Context” (PDF, under review at a NeurIPS 2026 workshop). The paper asks where the ability appears on a parameter ladder. This asks the question a buyer actually has: which one do I deploy, and what does it cost me all year.
Ten small models, 400 pre-registered questions, three runs each, by Alex Kwon. Verdict in one line: one model clears the bar and it is the one you cannot own, and all of them lean harder on your retrieval than on their own parameters.
Everything below is measured, not estimated, except one clearly-marked frontier price band. Data, scoring code and logs: pokemon-gen1-eval.
What we tested
The job
Answer a question by reading a reference document we put in the prompt, and give one specific answer. Same shape as looking something up in a policy, a spec, a fee schedule or a price list.
Why these models, and not the famous ones
Every model here is small enough to be cheap at volume, and all but one can run on hardware you own or rent by the hour. We are looking for the smallest thing that does the job, because this task runs thousands of times a day and nothing about it needs a frontier model.
What counts as correct
The exact right answer, nothing else. Close does not count. Every model ran three times, so we can see how much a score moves between runs rather than reporting one lucky attempt.
The control we ran
The same 400 questions with the document removed. A model that scores the same either way was not reading the document, and would score the same on a document that said something different.
How this maps to a real deployment
Your case has the same shape: someone asks a question, the answer is in a document you already have, and it has to be found and read correctly every time, thousands of times a day. So the number that matters is not how clever a model is in general. It is whether it reads the document you hand it, what it does when the wrong document turns up, and what it costs you to run all year.
One model clears the bar, and it is the one you cannot own
How often each model got the answer exactly right, with the reference document supplied.
train it
Bar is the score. The line through it is how much the score could move on a different 400 questions, so two bars whose lines overlap are not meaningfully different however different the numbers look. The dashed line is the 90% bar; models that do not clearly reach it are faded. Cost is dollars per 1,000 correct answers, so a cheap model that is often wrong does not look cheap.
Every model that works is leaning hard on the document you give it
The same questions with the document taken away. The gap between the two bars is what your search and retrieval is worth.
The same 400 questions asked twice: once with the reference document in the prompt, once without it. The second bar is what the model already knew. A small drop is bad news, not good, because it means the model was never really reading what you gave it.
Even the cheapest frontier model would cost twenty times more
What each option costs to run for a year at 100,000 lookups a day. The top bar is a range because frontier models are not priced alike.
cheapest tier: $380,969
Which frontier model? It does not change the answer, which is why the top bar is a range rather than a number. Frontier pricing spans roughly cheapest frontier tier at $1.25 in / $10.00 out per million tokens; mid frontier tier at $3.00 in / $15.00 out per million tokens; top frontier tier at $15.00 in / $75.00 out per million tokens. Across that whole range the cost lands between $380,969 and $3,421,875 a year, against $18,407 for gpt-5-nano. Even the cheapest frontier tier is 21× more.
If your provider charges less than any of these, substitute their rate: the arithmetic is a prompt of about 2,750 tokens and roughly 700 tokens of answer, 100,000 times a day. For a frontier model to actually match gpt-5-nano here it would need to charge around $0.08 per million input tokens, which is nano-class pricing. The cheapest frontier tier is about 15× that.
Faded bars are models that do not clear the accuracy bar, so they are not cheaper alternatives to anything. No frontier model was tested: its band assumes it would match the best score measured here, which is generous to it, so the conclusion holds even if that assumption is wrong. Published list rates on the day we ran this, via one aggregator. Read the shape, not the figures.
What this says to do
- Do not put a frontier model on this task.Depending on which one, it would cost between $380,000 and $3.4m a year at this volume to answer a question that is already printed in the document you supplied. The best small model here does the same job for about $18,000. The cheapest frontier option is still 21 times more.
- Start with gpt-5-nano.It is the only model that clears 90%, and it is the cheapest thing on the list that does. Go with it unless one of the trade-offs below is a dealbreaker.
- Spend your budget on finding the right document, not on a better model.Every model that works loses a quarter to a third of its score when the document is taken away. That drop is bigger than the gap between any two models on this list. If your search pulls the wrong page, the best model available answers confidently and wrongly.
- Know what you give up by picking gpt-5-nano.It is the one model here you cannot run on your own hardware, cannot train on your own data, and cannot pin to a version. If it gets something wrong in a way that matters to you, your only fix is to switch models.
- If you need to own it, qwen3-32b is the one, and budget to close the gap.It runs on a single graphics card you can buy, it stays inside your network, and you can train it on your own examples. It starts 13 points behind. Closing that is a project with a cost you can estimate, which is not true of any other option here.
- Do not use anything smaller than gemma-3-27b for this.Everything below it scores near what you would get by guessing. Cheap does not help when the answer is wrong.
Worth knowing before quoting any of this
- A short dark bar is bad news, not good. The four weakest models barely drop when we take the document away, because they were never really reading it.
- Prices are the published rates on the day we ran this, not quotes. Check them before you sign anything.
- The line through each bar is how much the score could move if we ran it again on a different 400 questions. Two bars whose lines overlap are not meaningfully different, however different the numbers look.
- This is one job. It tells you who reads a document well. It does not tell you who writes good email or who codes.
- You may pay less than the figures shown. Hosts charge different amounts for the identical prompt depending on whether they cache it, and one of these charged us a fraction of the rest. We have priced every model on the same prompt size so the comparison is between models, not hosts.
- We did not measure speed, on purpose. How fast an answer comes back depends on your provider, your region, how busy they are that hour and how you batch requests, so a time measured on our setup would not predict yours. What does carry over is how much each model writes to reach an answer, because every token costs time: here that ran from 63 tokens to 917, a 15x spread. Worth knowing before you test: the two models we recommend are also the two longest writers here, so expect them to be the slowest of the group, not the fastest. Once you have picked two finalists, time them on your own setup with your own traffic. That takes an afternoon and it is the only speed number worth having.
Pre-registered: the item set, the scoring rule and the predictions were hashed before any model was run, and the hashed files were never edited afterwards. Intervals are Wilson 95% over items, never items times epochs. Scorers were negative-tested against a mock that answers the key (must score 1.0), the key plus one (must score 0.0) and truncated output (must score 0.0 with a parse failure). Every host was pinned so a silent provider swap cannot move a number. Analysis and code: Alex Kwon. Full workings, the pre-registration, the findings ledger and every log: the pokemon-gen1-eval repository.
Cite: Kwon, A. (2026). Reference, Retrieve, Resolve: How Small a Language Model Can Read a Reference in Context. Under review, NeurIPS 2026 workshop. Paper and data: github.com/collapseindex/pokemon-gen1-eval (machine-readable: CITATION.cff in the repository). License: report CC BY 4.0; code Apache-2.0. Disclosure: AI-assisted implementation; the pre-registration, the controls and the interpretation are mine.
Alex Kwon · ask@collapseindex.org · more case studies · github.com/collapseindex