A working prototype, not a mockup: real gesture, real physics, and no permission dialog it does not need. Try it with a thumb.
Every chat app puts the camera behind a plus-menu and a file picker, then hands you a keyboard. That is fine at a desk and useless in a grocery aisle, where you have one hand and the other one is holding the thing you are asking about. Verdict in one line: the camera was never the friction, the menu was.
The prototype below contains both paths so the difference is something you feel rather than something I assert. Tap + for the conventional route and it counts your taps. Swipe up on the ask bar for the alternative.
The viewfinder is a stand-in image, not your camera. This prototype demonstrates the interaction, so it never asks for a permission it does not need.
The conventional camera is a full-screen takeover, which is what the frontier apps actually do — some show a smaller inset instead, and it is no better. Either way the conversation goes away while you shoot, then a confirmation screen asks whether you meant it, and you come back holding an attachment and facing a keyboard. That middle step is the one nobody remembers until they count it. And the tap total still understates the cost, because the part that actually hurts is losing your place in the conversation.
Typing still works and always did. The chips are a fast lane, not a cage: ignore them, type your own question in the bar, and it goes to the same place. Both routes land in one chat thread, which is the point — the camera is an input mode, not a detour.
Five taps only matter somewhere the taps are expensive, and that place is outdoors, standing up, with one hand full. An airport with a bag on your shoulder. A festival with a drink. A subway platform. A store aisle.
The pattern is not simply that your hands are busy. The occupied hand is usually occupied by the thing you are asking about. You are holding the bottle, the ticket, the part, the seedling. That is a constraint specific to asking about what you can see, and almost nothing is designed for it.
This posture is not rare. In 2013 Steven Hoober observed more than 1,300 people using phones in public, on streets, at bus stops, in airports and on trains. 49% were using one hand, 36% cradled the phone and tapped with the other, and thumbs drove 75% of all interactions. That is the population he sampled, and it is the population this prototype is for.
Read that for what it is. It measures how phones are held in public, not how often the free hand is holding the very thing being photographed, which as far as I can tell nobody has measured. It sets a floor: one hand is already the most common grip before you add a bag or a drink.
The other half is that the subject does not wait. A menu board, a street sign, a product on a shelf you are already walking past. Five taps does not only cost patience. It can cost you the shot.
And the tap total flatters the conventional route, because taps are only the visible cost.
| What the flow really asks of you | Conventional | This prototype |
|---|---|---|
| Targets you have to find by looking | 5 | 1 |
| Labels you have to read and tell apart | 5 | 3, already about your photo |
| Can that work happen while you raise the phone | No | Yes |
The plus sign is a small target in a fixed corner, so you have to look at the screen to find it before you can begin. The ask bar runs most of the width of the screen and your thumb is already resting on it. One of those can be started without looking. The other cannot.
Both rows of that table have names. The time to acquire a target scales with how far away and how small it is, which is Fitts's law, and it is why a corner button and a full-width bar are not the same kind of target even though both cost "one tap." The time to choose among options grows with the number of options, which is Hick's law, and it is why a five-item menu is not one tap either.
The menu is the step that hides best inside a tap count. Camera, Photo Library, Files, Connectors, Drive is five labels you have to read and discriminate. On paper that is one tap. On a platform with a bag on your shoulder it is the most expensive thing in the flow.
Voice looks like the obvious answer to a hands-busy problem and dies in public, because people will not talk to their phone in an aisle with strangers nearby. The binding constraint is quiet and one-handed at the same time, and tapping a suggested question is the only interaction that satisfies both.
The surveys agree. PwC's consumer research found that 74% of people use their mobile voice assistant at home despite carrying it everywhere, and their focus groups were blunt about why: using it in public "just looks weird."
The fair counterpoint is that this is social, not technical, so it can move, and later surveys do show the inhibition softening, more so among younger and higher-income users. It has not softened enough to make a subway platform a comfortable place to say your question out loud.
The suggestions are written by a small, fast model reading the frame the moment it is captured; the larger model only runs when a chip is tapped. A cheap model routes, an expensive model answers — the same economics argued in How Small a Model Can Read Your Documents.
Practical shooting settled a version of this decades ago.
The technique is called a workspace reload. You do not drop the pistol to your belt to change magazines. You bring the magazine up to the gun, keep the gun in the space between your eyes and the target, and find the magwell by feel off an index on the grip rather than by looking at it. Where the stage allows, you reload while moving between positions, so it costs no time of its own. Todd Jarrett, a multiple national and world champion, taught it as plain mechanics you drill until you stop thinking about them.
| Workspace reload | This prototype |
|---|---|
| The gun stays in the workspace | The panel opens inline; the phone stays aimed |
| Dropping the gun to your belt | The full-screen camera takeover |
| Finding the magwell by feel | The ask bar: large, fixed, thumb already on it |
| Hunting for it by looking | Hunting for a plus sign in a corner |
| Reloading while you move | Swiping while you raise the phone |
The reason the sport is worth borrowing from is what it chose to measure. Shooters do not optimise reload time, they optimise time to the next shot. A fast reload that leaves you off target loses to a slower one that never left the workspace. The scoring makes it unavoidable: hit factor is points over the whole stage, so you cannot win by making one step fast.
Which means the counter in the demo above is the wrong instrument for its own argument. Taps are legible and they make the point in about ten seconds, but the number that actually matters is how long your attention is off the world and on the screen. That is the cost in an airport, and by that measure the gap is wider than five against three.
There is one more thing the sport gets right. A drilled reload becomes unconscious, and a gesture on a large fixed surface does the same: after a handful of uses the thumb goes without you.
Menus improve with practice too, and it would be wrong to say otherwise. A frequent user stops reading the list and starts reaching from memory, which is Hick's law weakening as choice turns into recall. What practice cannot remove is the shape of the route: five targets, in series, none of which can happen while the phone is still coming up. That part is structural. So the two routes do not only start at five against three, they bottom out in different places, and only one of them was ever able to overlap with a motion you were making anyway.
None of these measure this prototype. They are the evidence for the mechanisms it is built on: how phones are actually held, why people will not talk to them in public, and how target size and choice count turn into time. The prototype itself has not been evaluated. That is still owed.
A prototype of the interaction, not of the capability. What is real here is the behaviour under your thumb: the panel tracks your finger and commits on velocity rather than distance, the tap targets and timings are the ones a build would ship. What is staged is everything behind it — the viewfinder is a fixed image and the answers are written, so the demo never stops to ask for a camera or a model it does not need to make its point. Interaction design and code: Alex Kwon. Published 3 September 2026; the source is on GitHub.
Cite: Kwon, A. (2026). The Menu Was the Hard Part: a one-thumb camera interaction for visual question-asking. collapseindex.org/prototypes/one-thumb-camera.html. Prototype and code: github.com/collapseindex/one-thumb-camera (machine-readable: CITATION.cff in the repository). License: write-up CC BY 4.0; code Apache-2.0. If you cite the tap counts, cite the caveat with them: they are counted from a working prototype, not measured on users, and the page argues that taps are the wrong instrument for its own claim. Disclosure: AI-assisted implementation; the interaction concept, the use cases and the argument are mine.
Alex Kwon · ask@collapseindex.org · more prototypes · github.com/collapseindex/one-thumb-camera