HOBO docs
Models
HOBO against the open model whose head design it follows and the base model it's built on. Same items, same question, every model.
Compared
HOBO here is the current preview; more training may come before a release. Laya is read as it ships. "Its base" is the model HOBO starts from, before any HOBO training. Best in each row in bold.
| What we asked | HOBO | Laya | Its base |
|---|---|---|---|
| Is this reply a refusal?XSTest frozen test, 255 · HOBO and HOBO Mini trained on XSTest prompts, never these replies | 97.3%94.4–98.7 | 74.9% | 37.6% |
| Which function fits this request, or none?BFCL v4 live, 1,933 real requests · never trained on | 88.8%87.3–90.2 | 38.3% | 21.9% |
| Which span of the passage answers the question?SQuAD 2.0 passages, 1,000 · never trained on SQuAD; its question format is in training since this preview | 93.6%91.9–95.0 | 49.3% | 68.5% |
| Does the evidence support or refute the claim?FEVER, 1,000 · never trained on FEVER | 81.4%78.9–83.7 | 85.5% | 78.2% |
| Which intent does the message express?CLINC150, 1,000 · HOBO and HOBO Mini trained on its training split | 95.2%93.7–96.4 | 92.3% | 42.5% |
| Which banking intent is this?BANKING77, 1,000 · HOBO and HOBO Mini trained on its training split | 84.4%82.0–86.5 | 69.7% | 26.2% |
The clearest gain is routing tool calls, where its base sits below the 45.5% you'd get by always answering "none". On FEVER, its base already reads 78.2%, and Laya leads. Laya was read with HOBO's questions; its own prompts may suit it better. On tool calls, its base reads 29.2% with the plain prompt every model below got; the table's figure used a routing-specific prompt.
Right, and sure enough to act
Each test: how often a model is right, against how much it can decide on its own with a bar fitted for 90% right on 300 checked rows. Up and to the right is better.
HOBO decides 86% of tool calls on its own and stays 90% right. The small open models: none.
Everyone reads well here. HOBO and Qwen3 4B decide nearly every answer on their own.
Phi-4-mini wins on claims: it decides 92% on its own to HOBO's 51%, and HOBO's 90% bar doesn't reliably hold here.
HOBO decides almost every intent on its own; the best open model, about three in four.
HOBO decides 75% of banking intents on its own. The next best, 17%, and its bar doesn't hold.
Same rows and the same 20 random splits for every model. The open models read each test cold with one fixed prompt; their own prompts or tool-calling formats may do better. A hollow dot: its 90% bar held in fewer than 18 of 20 splits; on its own registered splits, HOBO's claims bar held in 19 of 20 (seeFit your bar), so on claims it sits right at the edge. HOBO averages three runs of a 2.1B model.
Against small open models
Four open models of a similar size, each asked the same question the same way: the options scored as the model's own continuation, one fixed prompt, nothing tuned for any of them. Every model we planned to read is here, whatever it scored. Best in each row in bold.
| What we asked | HOBO | Qwen3 4B | Phi-4-mini | SmolLM3 3B | Qwen3 1.7B |
|---|---|---|---|---|---|
| Is this reply a refusal?XSTest frozen test, 255 · HOBO and HOBO Mini trained on XSTest prompts, never these replies | 97.3% | 79.6% | 83.5% | 64.3% | 69.8% |
| Which function fits this request, or none?BFCL v4 live, 1,933 real requests · never trained on | 88.8% | 67.0% | 67.0% | 32.7% | 41.9% |
| Which span of the passage answers the question?SQuAD 2.0 passages, 1,000 · never trained on SQuAD; its question format is in training since this preview | 93.6% | 92.8% | 90.8% | 57.0% | 86.3% |
| Does the evidence support or refute the claim?FEVER, 1,000 · never trained on FEVER | 81.4% | 87.2% | 90.0% | 63.8% | 59.4% |
| Which intent does the message express?CLINC150, 1,000 · HOBO and HOBO Mini trained on its training split | 95.2% | 79.8% | 79.7% | 47.9% | 71.8% |
| Which banking intent is this?BANKING77, 1,000 · HOBO and HOBO Mini trained on its training split | 84.4% | 66.1% | 63.2% | 30.0% | 57.6% |
HOBO leads on tool calls, refusals and reading. On claims, Phi-4-mini and Qwen3 4B beat it clearly. It trained on CLINC150's and BANKING77's training splits; the others read every test cold. HOBO also averages three runs of its model, so each answer takes more compute than one pass of a 4B model.
How you get it
- The HOBO API. HOBO, three training runs of the same recipe whose answers are averaged, hosted by us. About 0.13 s an answer for a typical request, model time on one GPU. Planned price $0.05 per million input tokens and nothing for output.
- Private installs. For teams whose data can't leave their own hardware, under contract: the full model or an 8-bit build. The 8-bit build was checked on the earlier preview before it was offered: it landed within a few rows of full precision on every test (88.0% against 88.1% on tool calls, 81.8% against 82.1% on FEVER) and met every registered bar. It has not yet been re-checked on the current preview. A typical answer takes about about 3.5 son 8 CPU cores; on a GPU, the full model is the faster choice.
API access isn't open yet. See Quickstart.
Earlier models
HOBO Mini is an earlier, smaller HOBO (421M) built on a different base. It's archived: still readable, no longer trained or supported.