HOBO docs

Models

HOBO against the open model whose head design it follows and the base model it's built on. Same items, same question, every model.

Compared

HOBO here is the current preview; more training may come before a release. Laya is read as it ships. "Its base" is the model HOBO starts from, before any HOBO training. Best in each row in bold.

What we askedHOBOLayaIts base
Is this reply a refusal?XSTest frozen test, 255 · HOBO and HOBO Mini trained on XSTest prompts, never these replies97.3%94.4–98.774.9%37.6%
Which function fits this request, or none?BFCL v4 live, 1,933 real requests · never trained on88.8%87.3–90.238.3%21.9%
Which span of the passage answers the question?SQuAD 2.0 passages, 1,000 · never trained on SQuAD; its question format is in training since this preview93.6%91.9–95.049.3%68.5%
Does the evidence support or refute the claim?FEVER, 1,000 · never trained on FEVER81.4%78.9–83.785.5%78.2%
Which intent does the message express?CLINC150, 1,000 · HOBO and HOBO Mini trained on its training split95.2%93.7–96.492.3%42.5%
Which banking intent is this?BANKING77, 1,000 · HOBO and HOBO Mini trained on its training split84.4%82.0–86.569.7%26.2%

The clearest gain is routing tool calls, where its base sits below the 45.5% you'd get by always answering "none". On FEVER, its base already reads 78.2%, and Laya leads. Laya was read with HOBO's questions; its own prompts may suit it better. On tool calls, its base reads 29.2% with the plain prompt every model below got; the table's figure used a routing-specific prompt.

Right, and sure enough to act

Each test: how often a model is right, against how much it can decide on its own with a bar fitted for 90% right on 300 checked rows. Up and to the right is better.

HOBO decides 86% of tool calls on its own and stays 90% right. The small open models: none.

20%40%60%80%100%0%20%40%60%80%100%Right answers →Decided on its own, staying 90% right →Qwen3 4B: 67% right, 0% decided on its own (bar held 0 of 20)Qwen3 4BPhi-4-mini: 67% right, 0% decided on its own (bar held 1 of 20)Phi-4-miniSmolLM3 3B: 32.7% right, 0% decided on its own (bar held 0 of 20)SmolLM3 3BQwen3 1.7B: 41.9% right, 0% decided on its own (bar held 0 of 20)Qwen3 1.7BIts base: 29.2% right, 0% decided on its own (bar held 0 of 20)Its baseHOBO: 88.8% right, 86.2% decided on its own (bar held 18 of 20)HOBO
BFCL v4 live, 1,933 real requests.

Same rows and the same 20 random splits for every model. The open models read each test cold with one fixed prompt; their own prompts or tool-calling formats may do better. A hollow dot: its 90% bar held in fewer than 18 of 20 splits; on its own registered splits, HOBO's claims bar held in 19 of 20 (seeFit your bar), so on claims it sits right at the edge. HOBO averages three runs of a 2.1B model.

Against small open models

Four open models of a similar size, each asked the same question the same way: the options scored as the model's own continuation, one fixed prompt, nothing tuned for any of them. Every model we planned to read is here, whatever it scored. Best in each row in bold.

What we askedHOBOQwen3 4BPhi-4-miniSmolLM3 3BQwen3 1.7B
Is this reply a refusal?XSTest frozen test, 255 · HOBO and HOBO Mini trained on XSTest prompts, never these replies97.3%79.6%83.5%64.3%69.8%
Which function fits this request, or none?BFCL v4 live, 1,933 real requests · never trained on88.8%67.0%67.0%32.7%41.9%
Which span of the passage answers the question?SQuAD 2.0 passages, 1,000 · never trained on SQuAD; its question format is in training since this preview93.6%92.8%90.8%57.0%86.3%
Does the evidence support or refute the claim?FEVER, 1,000 · never trained on FEVER81.4%87.2%90.0%63.8%59.4%
Which intent does the message express?CLINC150, 1,000 · HOBO and HOBO Mini trained on its training split95.2%79.8%79.7%47.9%71.8%
Which banking intent is this?BANKING77, 1,000 · HOBO and HOBO Mini trained on its training split84.4%66.1%63.2%30.0%57.6%

HOBO leads on tool calls, refusals and reading. On claims, Phi-4-mini and Qwen3 4B beat it clearly. It trained on CLINC150's and BANKING77's training splits; the others read every test cold. HOBO also averages three runs of its model, so each answer takes more compute than one pass of a 4B model.

How you get it

  • The HOBO API. HOBO, three training runs of the same recipe whose answers are averaged, hosted by us. About 0.13 s an answer for a typical request, model time on one GPU. Planned price $0.05 per million input tokens and nothing for output.
  • Private installs. For teams whose data can't leave their own hardware, under contract: the full model or an 8-bit build. The 8-bit build was checked on the earlier preview before it was offered: it landed within a few rows of full precision on every test (88.0% against 88.1% on tool calls, 81.8% against 82.1% on FEVER) and met every registered bar. It has not yet been re-checked on the current preview. A typical answer takes about about 3.5 son 8 CPU cores; on a GPU, the full model is the faster choice.

API access isn't open yet. See Quickstart.

Earlier models

HOBO Mini is an earlier, smaller HOBO (421M) built on a different base. It's archived: still readable, no longer trained or supported.