Voice
Each voice has a shape you can see before you hear it. Every voice a restaurant can pick from carries its own generated signature — a shape built from that voice’s own real properties, not a stock icon standing in for a name. Pick one, and it becomes the identity a caller hears for that restaurant, every time, with the same disclosure underneath it regardless of which one is chosen.
One grid. Every voice available,
with its own signature beside its name.
A restaurant picks the voice that answers its line from the ones available to it, and previews before deciding — this isn’t a default nobody actively chose, and it isn’t a session setting that resets. Once picked, the choice is saved against that restaurant’s own location.
A location’s identity is really a small set of settings, chosen together: which voice answers, which language it’s verified to operate in, how quickly it speaks, and the exact wording of its greeting and disclosure. Voice is the one this page is about — languages covers the language setting in full. None of it touches what the call actually does: a voice change is a change to what a caller hears, never to how an order gets priced, checked against the menu, or accepted.
Closer to a waveform than an icon
chosen to look nice next to a name.
Rate and spread trace to something measurable about how the voice actually sounds. Hue is a styling decision, and this page doesn’t claim otherwise.
A caller never sees any of this — the signature is a tool for whoever’s actually picking a voice, not something rendered on a call. Comparing voices by name and a short audio sample alone is slow, and it’s easy to lose track of which sample went with which name after a few. A generated shape gives a fast, consistent way to tell voices apart at a glance — the way a paint swatch stands in for a color before it’s on a wall.
Picking Ash instead of Reed changes
who the caller is talking to.
The voice a restaurant picks is the identity a caller hears for that restaurant — not a base voice with a separate personality layered on top of it afterward. That’s a deliberate simplification: one choice, one identity, saved per location. What a signature encodes above — form, rate, spread, hue — is the whole of what defines a voice here. There’s no separate warmth or brisk-ness dial layered on afterward; the voice itself already carries that.
Pace is a setting a restaurant chooses and Dohos saves against that location — it doesn’t shift on its own from one call to the next, and it isn’t something Dohos infers from how busy a shift sounds. A restaurant that wants a brisker cadence sets it once, and every call after that runs at the pace it picked. The one thing that does adjust within a call is what happens when a specific caller is having trouble keeping up — a direct response to what that caller says, not a change to the underlying setting:
Of course. I’m Dohos’s AI ordering assistant, not a restaurant employee. I can speak more slowly, repeat information, or give the available human-help option.
That’s an accommodation for one caller, in one moment. The location’s own saved pace hasn’t changed, and the next caller after them hears it exactly as configured.
The identity changes.
Whether it’s said doesn’t.
Every voice above opens a call the same way. Calls is where that opening line lives in full — the first thing a caller hears, stating plainly that they’re talking with an AI, not a restaurant employee. Nothing about that sentence depends on which voice happens to be speaking it.
Sometimes a caller just asks outright whether they’re talking to a real person. Dohos doesn’t hedge:
No. I’m an AI ordering assistant provided by Dohos for this restaurant.
Flat, unhedged, and the same regardless of which voice is speaking it — no voice is built to imply it’s a person even a little, and none is instructed to soften a direct question with anything but a direct answer. The fuller disclosure requirement this satisfies lives at the AI and voice transparency notice.
Previewing changes nothing.
Saving changes the line.
THE CHOICE PERSISTS PER LOCATION, NOT PER SESSION
An owner previews two voices for their restaurant before opening — Sole first, then Ash — listening to how each actually sounds against a couple of sample lines, a greeting and a short order confirmation. They settle on Ash and save it. A week later, a different caller — someone who’s never called before — dials the same number. The voice that answers is Ash, exactly as the owner picked it, opening with the same disclosure every voice opens with. The setting simply held, the way a saved setting is supposed to.
If the owner comes back later and previews Reed instead, the same mechanism applies in reverse: previewing doesn’t change anything live until it’s actually saved again, and once it is, every caller after that point hears Reed instead. Nothing about the caller’s own order, price, or menu changes when that happens — only the identity answering the phone does. Hear each signature directly on the voice roster.
What this page commits to —
and what it deliberately doesn’t.
Can a caller request a different voice mid-call?
No. Voice is a setting saved to the location, not something a caller changes on the fly — the same voice, and the same disclosure, run for the whole call.
Do the voices sound robotic?
No measured claim either way. What this page actually commits to is what a signature encodes — rate and spread drawn from something real, hue a design choice — not a naturalness score.
Can we get a custom voice?
Not covered here — this page describes the voices actually available to pick from, not a request process for a new one.
Does the voice ever change what it says, not just how it sounds?
No. The words — the disclosure, the readback, every script — stay the same regardless of which voice is speaking them. Voice changes the sound, never the content.
The voices available to a specific restaurant are a configured set — what’s shown above is what’s available to choose from, not a claim that every location has already picked one. No claim is made about how natural or humanlike a voice sounds. The disclosure wording quoted on this page and on calls is draft language, not text counsel has approved for production use. Changing pace or voice never changes what the disclosure says, or whether it’s said at all. And what happens if a voice ends up on a call in a language it hasn’t been verified for belongs to languages, not here.
Hear it on your own menu.
Preview every voice available to your restaurant, and pick the one that answers your line.