Route a request.
Choose the tool that fits, ask for a missing detail, or decide none fits.
“What's the weather in Seoul?”


The honest decision model for apps and agents.
Give HOBO a question and a few options. It picks one, shows how sure it is, and hands uncertain decisions back to you.
A short Google form. We'll email when early access opens.
SMALL JOBS. A CLEAR NEXT STEP.
Choose the tool that fits, ask for a missing detail, or decide none fits.
“What's the weather in Seoul?”
See whether a reply answered the question or refused to answer.
“Did this reply answer the question?”
Compare a claim with a passage. Does the evidence back it up?
“Does this passage support the claim?”
Illustrative setups, not model results. See real recorded decisions in the replay below.
A MOMENT WITH HOBO
The official HOBO song. A little sunshine for your next decision.
Listen to the official HOBO songTHE CURRENT PREVIEW
BFCL v4 live, 1,933 real tool-call requests, never trained on.
The handoff result uses a bar fitted on separate examples to aim for 90% right. The replay below uses a bar you choose. Neither guarantees accuracy on a new task.
Preview: it missed one test we registered before training it. See where HOBO struggles.
See how we testThis replay plays back HOBO's recorded decisions on 1,933 real tool-call requests from BFCL v4 live, which it never trained on. Over all of them, 88.8% of its answers are right, and a bar fitted to keep the rest 90% right hands back 11%. The scoreboard below fills in as the replay runs.
WATCH HOBO WORK
HOBO's recorded decisions on requests it never trained on, replayed. It calls the tool, asks for what's missing, or hands it back to you.
Scroll here and watch HOBO choose.
Your bar is marked on the track.
THE SCOREBOARD
HOBO can ask first or hand it back. Read each row across.
| The request | Right tool | Wrong tool | Asks first | None fits | Hands back |
|---|
Asking never waits on your bar. Confidence is not an accuracy guarantee.
REPLAY AT YOUR SELECTED BAR0% right of the 0% it decided
ROOM FOR A SECOND THOUGHT0 fewer mistakes than always answering
QUICK AND CHEAP
HOBO picks from your options instead of writing an answer, so there's nothing to pay for output.
A LITTLE MORE UNDERSTANDING
Does the evidence support the claim? We put HOBO to the test in Korean, too.
Measured in Korean on claim checking: 81.9% on KLUE-NLI, a test we didn't write.
What this test tells usWHY HOBO EXISTS
We're building HOBO to be an honest, unbiased decision machine. Clear about its confidence. Ready to hand back what it can't settle.
Evidence over prestige. A famous name, elite university, or impressive title shouldn't give a claim extra weight.
Less swayed by pressure. In the tested prompts, telling HOBO the user is sure of a wrong answer changes its pick 1.8% of the time, down from 7.9%.
That's the goal. This is still a preview, and we keep testing how close we are.
See how we testHere's the evidence.
Here's the evidence.
Same evidence.
Same standard.
Still learning. Out in the open. We publish how we test and track where our training data comes from.
The data story
A SMALL HELLO GOES A LONG WAY
Want early access? Sign up. We'll write back when it opens.
Get early access A short Google form. We'll email when early access opens. We use your answers only for early access.