# Set up a voice agent

The order the voice settings depend on each other in, why that order matters, and how to test before a real caller does.

Source: https://docs.omazy.ai/how-to/voice/setup/

import Figure from '../../../../components/Figure.astro'
import { Steps, Aside } from '@astrojs/starlight/components'

Voice is added to an agent that already works. Get the agent answering well in
chat first: the brief, the knowledge, the catalog. A voice agent that says the
wrong thing just says it out loud, faster.

## The order things depend on each other

Voice settings form a chain. Each one is meaningless until the one before it is
decided, and the last step is the one people skip.

<Figure
  label="Voice settings as a dependency chain ending at the call"
  caption="Warming is the step that has no visible result until you skip it. The phrases are already correct without it; they are just slow the first time each caller hears them."
>
<svg viewBox="0 0 700 190" xmlns="http://www.w3.org/2000/svg">
  <text x="0" y="14" class="d-eyebrow">DECIDE IN THIS ORDER</text>

  <rect x="0" y="28" width="126" height="46" rx="8" class="d-box" />
  <text x="12" y="48" class="d-label">Agent</text>
  <text x="12" y="64" class="d-sub">brief, knowledge</text>

  <path d="M130 51 L137 51" class="d-arrow" />
  <polygon points="143,51 136,47.5 136,54.5" class="d-arrow-head" />

  <rect x="143" y="28" width="126" height="46" rx="8" class="d-box" />
  <text x="155" y="48" class="d-label">Languages</text>
  <text x="155" y="64" class="d-sub">and their order</text>

  <path d="M273 51 L280 51" class="d-arrow" />
  <polygon points="286,51 279,47.5 279,54.5" class="d-arrow-head" />

  <rect x="286" y="28" width="126" height="46" rx="8" class="d-box" />
  <text x="298" y="48" class="d-label">Voice</text>
  <text x="298" y="64" class="d-sub">one per language</text>

  <path d="M416 51 L423 51" class="d-arrow" />
  <polygon points="429,51 422,47.5 422,54.5" class="d-arrow-head" />

  <rect x="429" y="28" width="126" height="46" rx="8" class="d-box" />
  <text x="441" y="48" class="d-label">Phrases</text>
  <text x="441" y="64" class="d-sub">per language</text>

  <path d="M559 51 L566 51" class="d-arrow" />
  <polygon points="572,51 565,47.5 565,54.5" class="d-arrow-head" />

  <rect x="572" y="28" width="126" height="46" rx="8" class="d-box-accent d-pulse" />
  <text x="584" y="48" class="d-label">Warm</text>
  <text x="584" y="64" class="d-sub">record them</text>

  <path d="M349 82 L349 104" class="d-arrow" />
  <polygon points="349,110 345.5,103 352.5,103" class="d-arrow-head" />

  <rect x="0" y="112" width="698" height="44" rx="8" class="d-box" />
  <text x="16" y="132" class="d-label">Incoming call</text>
  <text x="16" y="148" class="d-sub">resolves the whole chain, per call, in this order</text>
</svg>
</Figure>

Change something high in the chain and everything below it is affected. Pick a
new voice and the recorded phrases belong to the old one. That is the whole
reason [changing a voice](/how-to/voice/changing-voice/) has its own page.

## Setting it up

<Steps>

1. **Turn voice on for the app.**

   Go to **Agent, Saved Responses** and open the **Voice** tab. Voice starts
   off. While it is off you can author everything on this page and no caller
   will meet any of it, which is the right way to prepare a change.

2. **Choose the languages, in order.**

   The first language is the one the greeting leads with. Put the language your
   callers actually open in first, not the one your business writes in.

   Listening and answering are separate. An agent can answer only in Bengali
   and still understand a caller who says "my TIN number" in English, and it
   should: callers code-switch constantly, and a recognizer restricted to one
   language turns the switched words into nonsense.

3. **Pick a voice for each language.**

   One voice per language, not one voice overall. A single global voice reads
   English in the Bengali voice and the reverse, and both sound wrong.

   Listen before you commit. A voice that is merely a bad reader is not an
   error anywhere; it produces perfectly valid audio of the wrong impression.

4. **Write the spoken phrases.**

   At minimum a greeting. In practice a greeting plus three or four
   acknowledgements. See [Spoken phrases](/how-to/voice/phrases/) for what each kind
   is for and how long each should be.

5. **Warm them.**

   Warming records each phrase in the chosen voice ahead of time so the caller
   never waits for it. Without it every phrase is recorded on demand the first
   time it is needed, which lands on a real caller.

6. **Enable it, then attach a number.**

   Enabling is the switch that makes the call path read everything above
   instead of the platform defaults. Then point a phone number at the agent
   under **Voice, Numbers**.

7. **Call it yourself.**

   Not optional. Config that looks correct and a call that sounds correct are
   different claims, and only one of them is the product.

</Steps>

<Aside type="caution" title="Nothing above is live until you enable it">
An app with voice off keeps running on the platform defaults no matter what you
have authored. This is deliberate: it lets you stage a whole voice
configuration, review it, and switch it on in one move rather than exposing
callers to a half-finished line.
</Aside>

## Testing without a phone number

The **web dialer** is a browser-based call to the same agent, over the same
path a real call takes. It is the fastest way to hear a change.

It asks for microphone access before it shows a call button. That order is
intentional. A dial button that only discovers the microphone is blocked after
you press it produces a call that connects to silence, and the caller blames
the agent.

<Aside type="tip" title="Demo links are shareable, and that is the risk">
A dialer link works for anyone who opens it, and every call it places is a real
call against your agent. Treat one like a password: send it to a named person,
not a channel, and expect it to expire.
</Aside>

## What to check before you call it done

- The greeting plays immediately, with no pause after the line connects.
- The greeting is in the voice you picked, not a default.
- You can interrupt the agent mid-sentence and it stops.
- Ask something in each language you enabled and confirm the reply comes back
  in the language you expect.
- Say nothing for thirty seconds and confirm the line does something sensible.
