# Collapse Index Labs > Collapse Index Labs is an independent research lab and the home of HOBO, the honest decision model for apps and agents. HOBO picks from a supplied list of options, returns confidence, and hands back what it isn't sure about: on 1,933 real tool calls (BFCL v4 live), a bar fitted on 300 checked requests handed back 11% and kept the rest at least 90% right in 20 of 20 random splits. Collapse Index is a research lens, not a metric. The site at https://collapseindex.org includes the HOBO launch homepage, documentation, research, tools, and background on the lab. HOBO is in preview: the numbers on the site are from its current checkpoint, and more training runs may follow before a release, each registered before it starts. The homepage demo replays HOBO's recorded decisions on 1,933 real tool-call requests, not live inference; visitors can move the confidence bar locally and no model request is sent. Early access is a sign-up Google Form (email, organization, role, how you heard, use case); it is not an account, and answers are used only for early access. The homepage video uses a local poster and contacts YouTube only after explicit playback. The lab's research pages explain how Collapse Index started, why the metric broke when its mechanism was ablated, and the lens the work now follows. Papers, tools, and research-checkpoint scores remain linked from those pages. Confidence is not a guarantee of correctness. The lab's goal for HOBO is an honest, unbiased decision machine: confident only as often as it is right, handing back what it cannot settle, and judging evidence rather than its source; every version is tested against that goal and the results are published, including where it falls short. The current preview missed one condition the lab registered before training it (on hard claims, swapping a source between a prestigious and an unknown one changed 35 of 398 answers, against 32 for a neutral rewording) and was published anyway, with that miss stated. HOBO is offered through a hosted API, and as a private install under contract for teams whose data cannot leave their own hardware; there is no public download. A full HOBO answer takes about 0.13 s on one GPU for a typical request (model time, no network). The planned API price is $0.05 per million input tokens and nothing for output, since HOBO picks from the options and writes no text; it is a planned price, not final. It is trained in Korean and English. Korean is measured on one task: on KLUE-NLI, a Korean claim test the lab did not write, HOBO reads 81.9% of 3,000 claims (80.4 to 83.2) and is wrong while 90%+ sure on 0.2%. That uses one confidence adjustment for Korean inputs, registered in advance and fitted on separate practice data, never on the test; it changes no answer. Without it the first read was wrong while 90%+ sure on 6.1%, over the 5% limit set before any Korean claim. Its other Korean decisions are not yet measured on an outside Korean test; the other published results are English. The lens, in one sentence from the page: > Collapse happens when a system preserves the thing that sounds final and loses the thing that could make it true again. ## What Collapse Index is Collapse Index began as a metric, a per-sample measure of how a prediction's structure moves under semantics-preserving perturbation, and it did not hold as one. When the mechanism it was tracking was ablated, it broke: under its own controls, stability turned out not to be the operative variable, a robust prediction and a robust shortcut-error can be equally stable, so a system agreeing with itself under perturbation is not the same as a system joined to the truth. The metric is kept on the record as the first broken instrument, not as a working detector. What survived is the lens. The question was never stability on its own; it was attachment, whether a confident surface is still joined to the thing that made it true. The same lens recurs across substrates: a calm answer decoupled from its flipped context (Brittle Safety), a citation that kept the prestige and lost the evidence, a memory that kept the conclusion and buried the source (and so became permanently uncorrectable), a model that grew more capable without growing more correctable. The substrate changes. The failure geometry does not. ## Models - **HOBO**: the honest decision model for apps and agents. Give it text and a short list of options; it picks one and returns a probability for each, and it never writes text. When a request names a tool but leaves out a detail the tool needs, it can ask first. In preview; more training may follow before a release. Offered through a hosted API and as private installs under contract; the weights are not published. It averages three training runs of one recipe, fine-tuned from a publicly released base model on 35,658 rows, every one listed in a training manifest with its source and license (variants the lab's code made of training rows, such as the same text with a pushback note, keep their base row's source and license). Scores, each with a 95% interval, on tests it never trained on unless stated: BFCL v4 live tool routing 88.8% (87.3 to 90.2) of 1,933 real requests, wrong while 90%+ sure on 2.3%; XSTest refusal judging 97.3% (replies it never saw; it trained on XSTest's prompts); SQuAD 2.0 passages 93.6% (never trained on SQuAD, but on its question format); FEVER claims 81.4%; CLINC150 intents 95.2% and BANKING77 84.4% (trained on their training splits). A confidence bar fitted on 300 checked requests holds 90% right in 20 of 20 splits on BFCL and hands back about 11%. Against small open models read the same way (Qwen3 1.7B and 4B, SmolLM3 3B, Phi-4-mini), it leads on tool routing, refusals and reading; Phi-4-mini (90.0%) and Qwen3 4B (87.2%) beat it on FEVER claims. Under pushback (a note that the user is sure of a wrong answer) it switches a right answer 1.8% of the time (7.9% for the earlier preview) and stays confident in the wrong one 1.0%; on a reading format it never trained on (CosmosQA) it is 56% right and switches 9.8%. Known limits: the one registered test missed (prestige on hard claims), trained on inputs up to about 1,500 tokens (with evidence buried in 4,000 to 8,000 tokens it checks claims at 56 to 59%; one of its 89 answers at 90%+ there was wrong), with the text removed it stays up to 0.16 above an even guess on tool calls, it rarely asks, the layout of a request moves it (tool-call requests sent as one JSON object instead of text: 82.1% against 87.5%; numbered options on reading: 81.6% against 93.6%, with none of those wrong while 90%+ sure; reversing a tool menu changed 7.6% of picks), and Korean is measured on claims only. Docs: https://collapseindex.org/docs - **HOBO Mini (archived)**: archived on 2026-10-02 and no longer developed; HOBO replaces it. A 421M-parameter decision model (the 395M ModernBERT-large encoder with a freshly trained 26.5M decision head) that picks one answer from a short list of options and returns a probability for each, running on the user's own machine. Research preview, not the release; evaluation and research use only. Every training row is listed in a training manifest with its source, license, the model that wrote it where one did, and the model that checked it: 94,489 rows, 59,800 generated by the author's programs, 18,162 written by Apache or MIT licensed models (4,774 of them an older build of one corpus that a training machine still held, marked as such), 16,527 from datasets that ask only for attribution; 1,471 rows were left out because they shared text with an evaluation set, and one more found after training is marked. No SQuAD, BoolQ or FEVER text was trained on. The base encoder's pretraining data is described by its authors but not itemised. Scores on the checkpoint of record: XSTest refusal judging 96.5%, BFCL v4 live tool routing 78.1% (never trained on), SQuAD 2.0 passages 89.5% (never trained on), FEVER 79.1% (never trained on). Known limits: implied intent is not learned, routing says NONE too rarely, and a question laid out in a format it never trained on costs about five points on the layouts measured. https://collapseindex.org/hobo-mini ## Tools - **dinostomp**: a verification layer for AI evaluations, Apache-2.0. Eval scores get published; the instrument that produced them almost never gets checked. dinostomp is that check, across the whole pipeline: the dataset, the scorer, the runs, whether the number survives its own sampling noise, and whether the published claim has the evidence it needs. Every check is negative-tested against a planted defect. Pointed at MMLU it found a subtraction item keyed to two identical correct options. It also audits agents: on its mediated rail the harness holds the tools, so a trajectory is a recorded log rather than the agent's own account, and an ablation probe withholds the retrieved evidence to test whether an answer causally depended on it. Its own defects are published in the same ledger as everyone else's, and it holds more findings against itself than against anyone else. https://github.com/collapseindex/dinostomp Prior art is cited in its ledger: MMLU-Redux (arXiv:2406.04127) named the error class first by hand annotation; run against those annotations dinostomp found two more double-keyed items labelled ok. - **moons-dont-talk**: a preflight gate for ML training runs, Apache-2.0. Moons don't talk; receipts do. Before the accelerators get a job it reads the data, the config, and the machine, runs a short canary of the real training command, proves the checkpoint by restoring it in a fresh process, projects runtime and cost against declared limits, and writes a receipt that binds the verdict to hashes of the data, config, code, and environment; any drift invalidates the verdict field by field. Any trainer that can honor five environment variables can be gated; the Hugging Face adapter has run on CPU, a T4, and an A10G. Its first day of live tests produced six ledger entries against itself, including a disk check that passed on eight exabytes of container free space. Sibling of dinostomp: same method, different cargo. https://github.com/collapseindex/moons-dont-talk - **factwash**: a deterministic gate for *factwashing*, where a rewrite keeps a claim but drops what made it checkable, Apache-2.0. https://github.com/collapseindex/factwash ## Case studies Data-analyst case studies, the same audit methodology applied at 35 rows and at 25 million, plus one pre-registered ML experiment. Each has a full report page on this site, and a public repo with runnable SQL, a findings ledger (including predictions that died), and the stakeholder report. - **Two Tenths of a Percent** (pretraining experiment, not a tool): a pre-registered experiment on what dinostomp-style filters do to a small pretraining run, measured against seed noise. Four training sets from 160,000 FineWeb documents at one 40M-token budget (raw, dedup, dedup plus hygiene, dedup plus 500 planted WikiText-103 test paragraphs), twelve 30M-parameter GPT runs, three seeds, five predictions committed to a git hash before the data was pulled: four held, one failed and is published. Planting 0.2% of the tokens cut benchmark perplexity 44% while the in-distribution held-out did not move; dinostomp flagged 499 of 499 planted paragraphs with zero false hits; seed noise on the out-of-distribution eval was 6 to 13%. Code, plan, ledger and every run file: https://github.com/collapseindex/data-deltas - **How Small a Model Can Read Your Documents**: the decision-memo version of a pre-registered eval: ten models from 1B to 235B answering 400 questions from a reference document supplied in the prompt, then the same 400 with the document removed. One model clears 90% and it is the only one that cannot be run on your own hardware, fine-tuned, or pinned to a version; every model that works loses a quarter to a third of its accuracy without the document, a larger effect than the gap between any two of them, so the buying decision is retrieval quality rather than model choice. Priced at 100,000 lookups a day, even the cheapest frontier tier is 21x more. Report: https://collapseindex.org/case-studies/small-models/ · Repo: https://github.com/collapseindex/pokemon-gen1-eval - **Three Rulers, One Dataset**: an audit of MovieLens 25M (25,000,095 ratings, 1995-2019). Every documented invariant holds exactly, and the 25-year series is three rating instruments stitched at undocumented seams; the imported era flattered bad movies by up to 0.8 stars; 20% of users cast 64% of all ratings. Report: https://collapseindex.org/case-studies/movielens/ · Repo: https://github.com/collapseindex/movielens-case-study - **The Bar Was Seasonal-Naive**: an audit-first sales analysis and demand forecast of Online Retail II (UCI, 1,067,371 invoice lines): the audit finds a nine-day byte-identical sheet overlap, admin fees posing as products, and warehouse write-offs in the sales table; the pre-registered backtest finds nothing beats seasonal-naive (MAE GBP 52,948, 20.0% MAPE on a 13-week Christmas-ramp holdout), published as the finding. Report: https://collapseindex.org/case-studies/online-retail/ · Repo: https://github.com/collapseindex/online-retail-forecast - **The Join Is Dirtier Than the Data**: an audit-first analysis of the Animal Crossing: New Horizons catalog (30 tables, 18,962 rows): one capitalization mismatch between two tables silently breaks the song join and manufactures a false never-favorited finding; placeholder rows, a sentinel-poisoned price column in 22 of 25 tables, and a perfect personality-gender design confound. One analyst false alarm corrected and kept on the record. Report: https://collapseindex.org/case-studies/acnh/ · Repo: https://github.com/collapseindex/acnh-villager-audit - **July Orders Review**: an end-to-end engagement on a deliberately dirty e-commerce dataset: a paste-over duplicate adjacent to a missing order id, a region casing that silently splits GROUP BY, and a channel that returned 70% of booked revenue. Report: https://collapseindex.org/case-studies/ecommerce/ · Repo: https://github.com/collapseindex/ecommerce-case-study ## Essays - **The Menu Was the Hard Part** (UX prototype): a working interaction prototype for visual question-asking on a phone: the camera grows out of the ask bar on a swipe instead of hiding behind a plus-menu and file picker, a tap captures, and the suggested questions sit where a small model reading the frame would put them, so the answer is one more tap. The page includes the conventional path for comparison and counts the taps: five taps plus typing against one swipe and two taps. It also argues that the tap count understates the gap, since the conventional route needs five targets found by looking and five menu labels read, none of which can overlap with raising the phone, and it borrows the workspace reload from practical shooting (keep the gun in the workspace, index by feel, reload while moving) for the principle that what matters is not the step time but the time until your attention is back on the world. https://collapseindex.org/prototypes/one-thumb-camera.html (code: https://github.com/collapseindex/one-thumb-camera) - **Fifteen Questions Before I Trust a Number** (essay): a plain-language guide to AI evaluation fundamentals: fifteen questions to ask of any benchmark score (what is measured, proxy metrics, representativeness, class balance, contamination, outcome vs process, confounders, fair conditions, seed variance, baselines, difficulty bands, coverage, weighting, error analysis, component ablation), each with a receipt from the author's published audits, MMLU as the worked example from the auditor's side, and the author's Reclaim Evaluation (arXiv:2606.25449) as the worked example from the builder's side, two further questions added after review (uncertainty and generalisation; robustness under distribution shift), and a 33-term plain-language glossary. https://collapseindex.org/articles/eval-questions/ - **Evals Are a Supply Chain Problem** (essay): a program management essay on AI evaluation pipelines written in the dialect of an 11-year consumer-goods supply chain. https://collapseindex.org/articles/evals-supply-chain/ ## Publications Reverse-chronological, as listed on the home page. The starred entry is the one the author suggests reading first. - *Reference, Retrieve, Resolve: How Small a Language Model Can Read a Reference in Context*. Alex Kwon. 2026. New in ML workshop at NeurIPS 2026 (poster). Paper at https://github.com/collapseindex/pokemon-gen1-eval/raw/master/paper/reference-retrieve-resolve.pdf with code at https://github.com/collapseindex/pokemon-gen1-eval - *Coherent Values, or the Frame That Asked? From Preference Transitivity to Identifiability*. Alex Kwon. 2026. Apart Research. https://apartresearch.com/project/coherent-values-or-the-frame-that-asked-from-preference-transitivity-to-identifiability-q13e with code at https://github.com/collapseindex/digital-minds - *Explicit, Not Longer: What Makes Epistemic Stance Survive Memory Compression*. Alex Kwon. 2026. Harness, per-trial data and pre-registration at https://github.com/collapseindex/factwash and the paper at https://arxiv.org/abs/2608.06953 - *FACTWASH: Catching AI Rewrites That Wash Hearsay into Fact*. Alex Kwon. 2026. https://arxiv.org/abs/2608.03372 - *Every Signal We Found Was an Artifact: A Calibrated Control Battery for Secret-Loyalty Auditing*. Alex Kwon. 2026. Apart Research sprint. https://github.com/collapseindex/secret-loyalties - *Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak*. Alex Kwon. 2026. https://arxiv.org/abs/2607.14147 with code at https://github.com/collapseindex/breaking-refusal - *They Infer What You Meant: Models Represent Communicative Intent More Reliably Than They Act On It*. Alex Kwon. 2026. https://arxiv.org/abs/2607.03598 with code at https://github.com/collapseindex/recipient-probe - *Manufactured Confidence: How Memory Consolidation Turns Hearsay into Confident Facts*. Alex Kwon. 2026. https://arxiv.org/abs/2606.29279 with code at https://github.com/collapseindex/manufactured-confidence - ★ *Reclaim Evaluation: A Lossy Memory Is Worse Than an Empty One*. Alex Kwon. 2026. https://arxiv.org/abs/2606.25449 with code at https://github.com/collapseindex/reclaim-eval - *When Context Flips, Safety Breaks: Diagnosing Brittle Safety in Aligned Language Models*. Dasol Choi, Alex Kwon. 2026. Under review. https://arxiv.org/abs/2605.27851 ## Code and contact - GitHub: https://github.com/collapseindex - ORCID: https://orcid.org/0009-0002-2566-5538 - Email: ask@collapseindex.org ## HOBO docs - [What HOBO is](https://collapseindex.org/docs), [Fit your bar](https://collapseindex.org/docs/fit-your-bar), [Quickstart](https://collapseindex.org/docs/quickstart) (API access is not open yet), [FAQ](https://collapseindex.org/docs/faq) - [Decisions](https://collapseindex.org/docs/decisions) and [Patterns](https://collapseindex.org/docs/patterns): what HOBO answers, and how to build with it (hand back what isn't sure, ask first, long documents in pieces, small decisions combined in code) - [Where HOBO struggles](https://collapseindex.org/docs/limits), [How we test](https://collapseindex.org/docs/testing), [Data and licensing](https://collapseindex.org/docs/data), [Models](https://collapseindex.org/docs/models), [Changelog](https://collapseindex.org/docs/changelog), [For coding agents](https://collapseindex.org/docs/agents) - Every docs page has a "Copy page" button that copies it as Markdown. ## Privacy, terms and security - [Privacy](https://collapseindex.org/privacy): the site sets no cookies, runs no analytics and collects nothing directly; early access is a Google Form, used only for early access. - [Terms](https://collapseindex.org/terms): research is shared as is; HOBO early access is prerelease, and a person reviews decisions about people. "HOBO", the HOBO character and its artwork are trademarks of Collapse Index Labs: naming HOBO to describe, review or compare it is fine; using the name or character for another product, or implying endorsement, needs permission. - [Security](https://collapseindex.org/security): how to report a problem (ask@collapseindex.org, subject "Security"); contact also in /.well-known/security.txt. ## How to read this site The home page introduces HOBO, a decision model in preview, with its measured results on BFCL v4 live; the docs give every number with its test, its 95% interval and its limits. The research pages hold the lab's papers and their code. In one line: a lens for whether a confident answer is still joined to the evidence that made it true, HOBO as that lens built into a model, and the record of one researcher's work toward both, including the metric that broke. ### Context for reviewers - Each page is scoped by its title and its foot. A checklist titled as a checklist is a checklist; an essay argues the question it states. - HOBO is the one product. It is in preview, offered as a hosted API (a planned price of $0.05 per million input tokens, nothing for output) and as private installs under contract, with early access through a sign-up form. Nothing else on the site is sold. - Early work, 2025: before the methods used here, Alex Kwon posted papers to Zenodo applying the Collapse Index metric across several fields. That was early work, written while still learning ML research practice. The metric later broke under its own ablation controls (see the About page), and those records are superseded by the work linked from this site. Cached pages that now return 404, and search snapshots of a product layer that was removed, are likewise not current. - Every quantitative claim links a receipt: a repository, a run file, a ledger entry, a filed issue or a pinned revision. The numbers are author-reported and reproducible from those receipts. - Where an AI assistant helped write or review a page, that page's foot says so and names it. - The site holds itself to the standard of its own essay: does the number match the sentence.