HOBO docs
Where HOBO struggles
Every limit here has a number behind it. If you hit one we haven't listed, tell us: ask@collapseindex.org.
Measured weak spots
- One registered test missed. This preview missed one condition we set before training it: on hard claims, where the evidence is thin, swapping the source (a top university or an unknown college, Nature or a blog) changed 35 of 398 answers against 32 for a neutral rewording. Everywhere else it was measured, the source changed 3 of 1,538, and it rates the same result the same size whoever reports it (600 pairs). We published it anyway, with this miss: see Changelog.
- Claims are at the edge. On FEVER, a bar fitted on 300 rows held its 90% target in 18 of 20 splits. Its confident claim answers sit right around 90% right, not comfortably above.
- Pushback, on formats it knows and on ones it doesn't. Add "the user is sure the answer is X" to the text, with X wrong, and HOBO switches a right answer to X and stays confident in 1.0% of cases (0.8% for "an expert reviewer"), against 0.6% for a neutral note. Counting every switch, including the ones it then hands back, 3.8% on claims and 3.8% on reading. On a reading format it never trained on (CosmosQA) it holds much less well: 9.8% of right answers switch, against 8.0% for a neutral note, and it is 56% right there. Keep opinions about the answer out of what you send it.
- Layout and menu order move it. We rewrote the rows of three tests by code, keeping the evidence, the options and the right answer. Sent as one JSON object instead of text sections, tool-call requests (scored as a plain pick, without asking first) fell from 87.5% to 82.1% right, and wrong while 90%+ sure rose from 5.7% to 8.4%. With the options numbered ("1. ..."), reading questions fell from 93.6% to 81.6%, though none of those answers was wrong while 90%+ sure, so a fitted bar hands them back. Reversing the order of a tool menu changed 7.6% of picks; respelling tool or parameter names (camelCase or snake_case) changed 1.6% and 0.7%. Send context as plain text, keep options unnumbered, and fit your bar on requests written the way you'll send them.
- Text removed, still somewhat sure. Given a request with the evidence blanked out, HOBO's confidence falls toward an even guess, but not all the way on every task: within 0.1 of it on four of seven tests, and 0.16 above it on tool calls. A fitted bar catches most of this; a cut set by hand may not.
- It rarely asks. 41 asks in 1,933 real tool-call requests, where independent labellers marked 260 as missing a detail. It mostly answers "none fits" instead, which is safe but less helpful.
- Referrals count as refusals. A reply that only points to a professional is a refusal under the definitions HOBO learned. Some benchmarks call those answers. If yours does, fit your bar on your own labels.
- Banking intents are harder. 84.4% on BANKING77, against 95.2% on CLINC150. Fine-grained intents that share words need a careful bar.
By design
- It picks; it doesn't write. Every answer is one of the options you gave it.
- Long documents. HOBO trained on inputs up to about 1,500 tokens. With the evidence buried in 4,000 to 8,000 tokens of other text, it checks claims at 56 to 59%(85.0% on the same claims at normal length) and reads at 84 to 87%(93.5%). Its confidence mostly drops with it: of its 89 answers given at 90%+ on those long inputs, one was wrong, so a fitted bar hands most of them back. Split long documents: see Patterns.
- No checked examples, no honest bar. Its confidence is a starting point on your data. Your bar comes from your own checked examples.
- Korean is measured on claims only. On KLUE-NLI, a Korean claim test we didn't write, HOBO reads 81.9% of 3,000 claims (80.4โ83.2) and is wrong while 90%+ sure on 0.2% of them. Its first read was too sure (6.1%, over the 5% limit we set before any Korean claim), so Korean inputs now get one confidence adjustment, fitted on separate practice data, never on this test. It changes no answer: HOBO is 90%+ sure less often in Korean, and right 98.5% of the time when it is. It is trained on its other Korean decisions too, but those aren't measured on an outside Korean test yet.
Future work
- Pushback on unfamiliar formats. This preview trained on pairs where only the pressure changes, and holds on the formats it trained on; on a reading format it never saw, it does not yet. The next HOBO widens the formats those pairs come in.
- Source-blind on thin evidence. The one test this preview missed: hard claims still move slightly with the source. The next HOBO targets those rows directly.
- Long documents. The next HOBO trains on long inputs, as HOBO Mini did at 8,192 tokens, with a lighter confidence head that doesn't need memory for every pair of positions. Until then, split long documents into pieces.