HOBO docs

Fit your bar

The bar is how sure HOBO must be before your software acts on its answer. You fit it once per task, on your own checked examples. Everything under the bar comes back to you.

Why your own data

HOBO's confidence is trained to track how often it's right, and on the tests we publish it does. But your data isn't our tests. The only way to know what "90% sure" means on your requests is to check some of them by hand and let the numbers set the bar. That's what makes the promise yours, not ours.

The steps

  1. Check a few hundred examples by hand. We use 300. Label the right answer for each.
  2. Pick a target. For example: the answers we act on must be right at least 90% of the time.
  3. Find the lowest bar that meets it. Run HOBO on your labelled examples and find the lowest confidence at which the answers above it still meet your target. We use a Wilson lower bound, not the raw rate, so a small sample can't pass by luck.
  4. Act on what clears the bar. Everything under it goes to a person, or to a bigger model. It's a hand-off, not a guess.
  5. Refit when your data changes. A bar fitted on one kind of request doesn't carry to another.

What it looks like on real data

We fitted a bar the same way on 300 randomly chosen requests, 20 times, and checked it on the rest of each test:

TestTarget heldDecided on its own
Tool calls (BFCL v4 live, 1,933)20 of 2088.8%
Refusal requests (Do-Not-Answer, three request types)20 of 2099.8%
Claims (FEVER, 1,000)18 of 2050.7%

FEVER is at the edge: the target held in 18 of 20 splits, the fewest a bar may hold and still count. On FEVER-style claims, HOBO's confident answers sit right around 90%, not comfortably above it. Fit your bar with that in mind, and on more rows if you can.

How many examples

More checked examples let the bar sit lower, so HOBO decides more on its own; they don't make HOBO more accurate. As a floor: to show a 95% target, you need at least 52 checked examples of each answer, all right. For most tasks, a few hundred checked examples is the sweet spot.

Asking doesn't wait on the bar

When HOBO decides a request names the right tool but leaves out a detail the tool needs, it asks first. Asking the user is something your system does by itself, so it never waits on the bar. See Patterns.