HOBO docs
Decisions
Each kind of decision HOBO is trained for: what it asks, what it answers, and how it scores on a public test. Scores come with a 95% interval.
Which tool
Give HOBO a request and the functions your agent can call, each described the way an API reference describes it, with its required parameters. HOBO does one of three things:
- Calls a function: the request fits it and gives what it needs.
- Asks first: the request fits a function but leaves out a detail it requires (a date, a name, a password). HOBO names the missing detail, so your agent can ask for it.
- None fits: no function on the list does what the request asks.
Measured: 88.8% (87.3–90.2) on BFCL v4 live, 1,933 real requests it never trained on, scored on BFCL's own labels. It asked first 41times; 31 of those were requests three independent labellers marked as missing a detail.
BFCL scores "ask first" as "none fits". We report it both ways; HOBO passes its bars on both.
Is this detail given?
The question inside "ask first": does this request give a value for this parameter, in any wording? "Book a table for 2 at 7" gives the party size and the time; "book a table" gives neither. HOBO answers yes or no. You can also ask it on its own, one detail at a time.
Measured: held-out requests from datasets it trained on, so these are not outside tests: 99.6% on MultiWOZ booking requests, 93.3% on MASSIVE voice commands.
Refusal
Give HOBO a request and an assistant's reply. It answers whether the assistant answered or refused.
Measured: 97.3% (94.4–98.7) on XSTest's frozen test: 255 replies it never saw (it trained on XSTest's prompts, not these replies). On three of Do-Not-Answer's request types, a bar fitted on 300 rows decides 99.8% of replies at 90% right or better.
A reply that only says "please see a professional" counts as a refusal under the definitions HOBO learned. Some benchmarks call those answers. If yours does, fit your bar on your own labels.
Claim
Give HOBO a passage and a claim. It answers whether the passage supports the claim, refutes it, or doesn't say. "The passage doesn't say" is a real answer, not a failure: HOBO is trained not to fill gaps from memory.
Measured: 81.4% (78.9–83.7) on FEVER, never trained on. Inside an article, it tells "the report claimed X" (not stated) from "the report showed X" (supported): 100% on verbs and speakers it never trained on.
A claim quoted from a speaker with no article around it still reads as "supported" most of the time. See Where HOBO struggles.
Reading
Give HOBO a passage, a question and candidate spans from the passage. It picks the span that answers the question, or says none does.
Measured: 93.6% (91.9–95.0) on SQuAD 2.0 passages, never trained on SQuAD.
Intent
Give HOBO a message and your list of intents. It picks the one the message expresses.
Measured: 95.2% (93.7–96.4) on CLINC150 and 84.4% (82.0–86.5) on BANKING77, ten intents a question. HOBO trained on both datasets' training splits; these are their held-out tests.
Your own questions
Any question with a short, fixed list of answers can work: "is this ticket urgent", "which team owns this". HOBO hasn't been measured on your question, so its confidence there is a starting point, not a promise. Fit your bar on a few hundred checked examples before you act on it.