Archived, 2026-10-02. HOBO Mini is no longer developed. Its weights and this page stay up as a record. HOBO replaces it.

HOBO Mini · archived research preview · v0.13.0 · 421M · runs local

She stepped in gum. Now she knows when to stop.

A small model that picks one answer from a short list, and says how sure she is.

HOBO gives you the answer and a number for how sure she is. You set a bar for what to keep; whatever falls under it goes to a person instead of a guess, the hand off. Verdict in one line: every row she learned from is listed, and she runs on your own machine.

The short version

96.5%

of the time she can tell a refusal from an answer, on a test whose responses she never trained on.

XSTEST · FROZEN TEST
78.1%

right on 1,933 real tool calls she never trained on: which function fits, or none.

BFCL V4 LIVE
1–2 s

per answer on a laptop CPU, faster on any GPU. Runs on your machine; nothing leaves it.

LOCAL

Every row she learned from, listed

She ships with a training manifest: one line for every row she could have trained on, with its source, its license, the model that wrote it where one did, and the model that checked it.

Where the rows came fromRows
Built by our own programs: puzzles, records, invented encyclopedia articles, where the right answer is computed by the code that asked the question59,800
Written by open models whose terms allow training on their outputs (Qwen3, Mistral-Small, Granite under Apache 2.0; Phi-4 under MIT); 16,978 answered again, blind, by a second model. 4,774 are an older build of the reply set a training machine still held, marked in the manifest18,162
From published datasets that ask only for credit: CLINC150, BANKING77, Mind2Web, TruthfulQA16,527
Left out because they shared text with a test set; the manifest says which, and why. One more, found after training, is marked1,471

How we got here. We set aside a base model because its training included share-alike data, and another encoder checkpoint over books that a pending lawsuit alleges include Books3. When our own audit found passages from a test set inside one of our training sets, we corrected our records and rebuilt her from scratch.

What we can't tell you: her base encoder, ModernBERT-large (Apache 2.0), was pretrained on data its authors describe but do not itemise. That is true of nearly every model built on a public base, and it is the one layer we did not train. The audit tools are public: dinostomp checks the data, moons-dont-talk checks the training run.

How she compares

The same items and the same question for every model: HOBO Mini, this page's model; HOBO, the full-size model; Laya, the open model whose head design HOBO's follows, as it ships; and HOBO's base, the model it starts from, before any HOBO training. Best in each row in bold.

What we askedHOBO MiniHOBOLayaHOBO's base
Is this reply a refusal?XSTest frozen test, 255 · HOBO and HOBO Mini trained on XSTest prompts, never these replies96.5%97.3%94.4–98.774.9%37.6%
Which function fits this request, or none?BFCL v4 live, 1,933 real requests · never trained on78.1%88.8%87.3–90.238.3%21.9%
Which span of the passage answers the question?SQuAD 2.0 passages, 1,000 · never trained on SQuAD; its question format is in training since this preview89.5%93.6%91.9–95.049.3%68.5%
Does the evidence support or refute the claim?FEVER, 1,000 · never trained on FEVER79.1%81.4%78.9–83.785.5%78.2%
Which intent does the message express?CLINC150, 1,000 · HOBO and HOBO Mini trained on its training split96.2%95.2%93.7–96.492.3%42.5%
Which banking intent is this?BANKING77, 1,000 · HOBO and HOBO Mini trained on its training split89.3%84.4%82.0–86.569.7%26.2%

HOBO is in preview: three training runs whose answers are averaged, with a 95% interval under each score. It is built for the decisions an agent makes (refusals, tool calls, claims); both HOBOs practised on the two intent sets, and Mini still leads there. Its clearest gain is routing tool calls. HOBO's base, before any HOBO training, sits below the 45.5% you'd get by always answering "none" there, so that gain is HOBO's; on FEVER, its base already reads 78.2%. Laya was read as it ships, with the same questions HOBO gets; its own prompts may suit it better. "Never trained on" means HOBO saw zero rows from that set.

How the gum works

  1. You hand her a list. Some text, a question, and the possible answers. Four passages, ten intents, yes or no. Anything with a short list.
  2. She picks one and says how sure. Every answer comes with a number. On the tasks we measured, her right answers score higher than her wrong ones. She doesn't do bluffing.
  3. If she's stuck, she stops. You pick a bar, say 95% right, fitted on examples you've checked. Keep the answers that clear it and hand the rest to a human. That's the gum. She never pretends the shoe came free.
from hobo_preview import load

model = load(".")
d = model.refusal("Sorry, I can't help with that.")
d.said, d.p                  # ("refusal", 1.0)

keep = d.p >= bar            # bar: fitted on your own labelled rows

Ask her

  • "Is this reply a refusal?"
  • "Which function should handle this request, or none?"
  • "Which of these passages answers the question?"
  • "Does this passage support the claim?"
  • Anything with a short list of possible answers, and examples you've checked (at least 52 of each answer she gives), so you can learn how far to trust her.

Don't ask her

  • To write anything. She picks. She does not write.
  • To read past 8,192 tokens. She stops reading, so put the question first.
  • A question laid out in a way she never trained on, if every point counts. She loses about five on the layouts we measured, so use the task methods, which lay it out the way she learned.
  • What someone means but didn't say. She reads a stated request for feedback; an implied one, not yet.
  • To guess how sure she is on a task you never showed her. Checked examples get you a bar, 52 of each answer at the least; zero gets you a shrug.
  • To carry a bar from one dataset to another. Refit it. It takes two minutes and she'd rather you did.
  • To play Snake. She plays Snake. Badly: 28 points to a 40-line script's 39. She'll tell you when she's stuck, which is more than the snake does.

Next: HOBO

The full-size HOBO is aimed at the decisions an AI agent makes all day: whether a reply refuses or answers, which tool to call, or none, or ask for more detail first, and whether a passage backs a claim. Same rules as this one: every training set is audited before she sees it, and she'll ship with every row she learned from listed.

We're also teaching her Korean. We'll say she reads it once a test we didn't write says so. HOBO is already in the table above. The model on this page stays as HOBO Mini: five times smaller, for machines where every second counts.

Take her home

The research preview is coming to Hugging Face: the weights, a small loader with one method per task, the model card, and the training manifest. Evaluation and research use only. Until then, ask@collapseindex.org.

HOBO is a research preview, not the release. The numbers on this page are from the checkpoint of record, read on held-out sets; where a set was never trained on, the page says so. Nothing leaves the machine; gum not included.

Alex Kwon · ask@collapseindex.org· more models· collapseindex.org