HOBO docs

FAQ

Short answers. Each one links to the page with the long version.

The basics

What is HOBO?

A decision model. You give it some text and a short list of options; it picks one and says how sure it is. When it isn't sure enough for your bar, the decision comes back to you.

Can it write text or chat?

No. Every answer is one of the options you gave it. That's what makes its confidence checkable.

Is it a safety filter?

No. Judging whether a reply is a refusal is one of the decisions it's trained for, alongside tool calls, claims, reading and intents. It's built for any decision with a fixed list of answers.

Can I download it?

No. You use HOBO through the HOBO API. Teams whose data can't leave their own hardware can get a private install on your own hardware, under contract.

Confidence and the bar

What does its confidence mean?

It's the probability HOBO puts on each option, trained to track how often it's right. We check that on tests it never trained on: on real tool-call requests it's wrong while 90% sure or more on 2.3% of them.

Why do I need my own checked examples?

Because your data isn't our tests. A few hundred checked examples tell you what "90% sure" means on your requests, and set the bar from that. See Fit your bar.

How many checked examples?

We use 300. More let the bar sit lower, so HOBO decides more on its own. To show a 95% target you need at least 52 of each answer, all right.

What happens to the answers under the bar?

You decide. A person, a review queue, or a bigger model. HOBO tells you which side of the line each decision is on; it doesn't decide where the line goes.

Why does it sometimes ask instead of answering?

When a request fits one of your tools but leaves out a detail the tool needs, guessing is wrong and refusing is unhelpful. So it names the missing detail for your agent to ask about. See Patterns.

How good it is

Where is it weakest?

Claims. On FEVER it's 81.4%, and a bar fitted for 90% there holds only just. Claims quoted from a speaker with no article around them mostly read as "supported". See Where HOBO struggles.

Does it read Korean?

For checking claims, yes, measured. On KLUE-NLI, a Korean claim test we didn't write, HOBO reads 81.9% of 3,000 claims (80.4โ€“83.2) and is wrong while 90%+ sure on 0.2% of them. Its first read was too sure (6.1%, over the 5% limit we set before any Korean claim), so Korean inputs now get one confidence adjustment, fitted on separate practice data, never on this test. It changes no answer: HOBO is 90%+ sure less often in Korean, and right 98.5% of the time when it is. It's trained in Korean on its other decisions too, but those aren't measured on an outside Korean test yet.

Can I check the numbers?

Every training run and its pass or fail rules are written down before it starts, and misses are published. Each number in these docs names its test and comes with a 95% interval. See How we test.

Getting it

When can I use it?

HOBO is in preview, and API access isn't open yet. Early access comes first: sign up to hear when it does. See Early access sign-up.

What does a private install need?

A GPU, or a CPU if speed matters less. It comes at full precision or as an 8-bit build, which scored within a few rows of the full model on the earlier preview (not yet re-checked on this one) and needs less memory. See Models.

What was it trained on?

35,658 rows, each listed with its source and license. No non-commercial or share-alike data, no outputs from closed models, and no test sets. See Data and licensing.

What happened to HOBO Mini?

It's archived. HOBO replaces it; its page stays up for reference. See HOBO Mini.