HOBO and his companion walk along the coast, looking toward each other.
HOBO/Preview
HOBO

The honest decision model for apps and agents.

A clear choice.
Or back to you.

Give HOBO a question and a few options. It picks one, shows how sure it is, and hands uncertain decisions back to you.

A short Google form. We'll email when early access opens.

LESS GUESSWORK. MORE SAY.A LITTLE HELP, EVERY DAY

SMALL JOBS. A CLEAR NEXT STEP.

What could HOBO help with?

Explore the uses

Route a request.

Choose the tool that fits, ask for a missing detail, or decide none fits.

EXAMPLE SETUP

“What's the weather in Seoul?”

WeatherCalendarNo tool

Check a reply.

See whether a reply answered the question or refused to answer.

EXAMPLE SETUP

“Did this reply answer the question?”

AnsweredRefused

Check a claim.

Compare a claim with a passage. Does the evidence back it up?

EXAMPLE SETUP

“Does this passage support the claim?”

SupportsDisagreesDoesn't say

Illustrative setups, not model results. See real recorded decisions in the replay below.

A MOMENT WITH HOBO

A little HOBO.A little sunshine.

The official HOBO song. A little sunshine for your next decision.

Listen to the official HOBO song

THE CURRENT PREVIEW

HOBO, on real requests.

Preview

BFCL v4 live, 1,933 real tool-call requests, never trained on.

Across all requestsRight answers
88.8%95% interval 87.3–90.2
Across all requestsConfident mistakes
2.3%wrong, while 90%+ sure
With a fitted handoff barDecisions handed back
11%to keep the rest 90% right

The handoff result uses a bar fitted on separate examples to aim for 90% right. The replay below uses a bar you choose. Neither guarantees accuracy on a new task.

Preview: it missed one test we registered before training it. See where HOBO struggles.

See how we test

This replay plays back HOBO's recorded decisions on 1,933 real tool-call requests from BFCL v4 live, which it never trained on. Over all of them, 88.8% of its answers are right, and a bar fitted to keep the rest 90% right hands back 11%. The scoreboard below fills in as the replay runs.

WATCH HOBO WORK

1,933 real requests.
One call each.

About this replay

HOBO's recorded decisions on requests it never trained on, replayed. It calls the tool, asks for what's missing, or hands it back to you.

Ready when you are0 / 1,933
HOBO's current requestOpen to follow the latest pick
ONE REQUEST AT A TIME Recorded replay

Scroll here and watch HOBO choose.

HOBO'S PICK
A little thought first.

THE SCOREBOARD

Same requests. Two ways.

Filling as we go

HOBO can ask first or hand it back. Read each row across.

Rows show what the request needs. Columns show the recorded action.
The requestRight
tool
Wrong
tool
Asks
first
None
fits
Hands
back
✓Right!OK, asking was better×Wrong↩Back to you
MISTAKES, SIDE BY SIDESame requests so far
HOBO
0
Always answers
0

Asking never waits on your bar. Confidence is not an accuracy guarantee.

REPLAY AT YOUR SELECTED BAR0% right of the 0% it decided

ROOM FOR A SECOND THOUGHT0 fewer mistakes than always answering

QUICK AND CHEAP

Small decisions should cost small change.

Planned API price

HOBO picks from your options instead of writing an answer, so there's nothing to pay for output.

Answer time
0.13 sper answer on one GPU, typical request
Input
$0.05per million tokens, about 1¢ per 1,000 decisions
Output
$0it picks, it doesn't write
How you get it

A LITTLE MORE UNDERSTANDING

Korean claims.
Tested, too.

Does the evidence support the claim? We put HOBO to the test in Korean, too.

Measured in Korean on claim checking: 81.9% on KLUE-NLI, a test we didn't write.

What this test tells us

WHY HOBO EXISTS

Honest and fair.
That's the mission.

We're building HOBO to be an honest, unbiased decision machine. Clear about its confidence. Ready to hand back what it can't settle.

Evidence over prestige. A famous name, elite university, or impressive title shouldn't give a claim extra weight.

Less swayed by pressure. In the tested prompts, telling HOBO the user is sure of a wrong answer changes its pick 1.8% of the time, down from 7.9%.

That's the goal. This is still a preview, and we keep testing how close we are.

See how we test
A WELL-KNOWN NAME

Here's the evidence.

A NAME YOU DON'T KNOW

Here's the evidence.

Same evidence.
Same standard.

The standard we're working toward.

Still learning. Out in the open. We publish how we test and track where our training data comes from.

The data story

A SMALL HELLO GOES A LONG WAY

Be here for
HOBO's first day.

Want early access? Sign up. We'll write back when it opens.

Get early access A short Google form. We'll email when early access opens. We use your answers only for early access.