Agents
Transcriber tab
The Transcriber tab picks the speech-to-text service that turns the caller's speech into text and decides when the caller has finished speaking.
Fields
All fields live under transcriber in the agent object. Some fields apply to one provider only.
| Field | Type | Default | Allowed | Provider | What it does |
|---|---|---|---|---|---|
provider | string | cartesia | cartesia, deepgram, sarvam | all | The speech-to-text service. |
model | string | ink-2 | ink-2 (Cartesia), nova-3 (Deepgram), saaras:v3-realtime or saaras:v4 (Sarvam) | all | The model. If it doesn’t match the provider, Vaakyo replaces it: ink-2 for Cartesia, nova-3 for Deepgram, saaras:v3-realtime for Sarvam. |
turn_detection | string | balanced | responsive, balanced, patient | Cartesia, Sarvam | How quickly the agent answers once the caller pauses. On Sarvam: a pause of 0.5 s, 0.7 s or 1.1 s. |
language | string | hi | Deepgram: hi, multi, en-IN, en. Sarvam: auto, hi-IN, en-IN, bn-IN, gu-IN, kn-IN, ml-IN, mr-IN, od-IN, pa-IN, ta-IN, te-IN | Deepgram, Sarvam | The language the caller speaks. auto makes Sarvam detect it. For Sarvam, hi, en and multi are accepted and saved as hi-IN, en-IN and auto. Cartesia ignores it and detects the language itself. |
endpointing_ms | integer | 300 | 10 to 3000 | Deepgram | Milliseconds of silence that end the caller’s turn. |
keyterms | array of strings | [] | up to 100 | all | Words the recogniser should expect, such as doctor or product names. Cartesia uses up to 100, Deepgram the first 50, Sarvam the first 50 on saaras:v4 only. |
fallbacks | array | [] | up to 2 | all | Backup transcribers, connected in order when the transcriber fails. See Fallback transcribers. |
Cartesia example
{
"transcriber": {
"provider": "cartesia",
"model": "ink-2",
"turn_detection": "patient",
"keyterms": ["Dr. Kulkarni", "Aundh", "physiotherapy"]
}
}
Deepgram example
{
"transcriber": {
"provider": "deepgram",
"model": "nova-3",
"language": "multi",
"endpointing_ms": 400,
"keyterms": ["Dr. Kulkarni"]
}
}
Sarvam example
{
"transcriber": {
"provider": "sarvam",
"model": "saaras:v3-realtime",
"language": "hi-IN",
"turn_detection": "balanced"
}
}
Choosing a provider
Cartesia ink-2 | Deepgram nova-3 | Sarvam saaras | |
|---|---|---|---|
| Language | Detects it on its own (en, hi, fr, ja, es) | You set it: hi, multi (switches between languages), en-IN, en | You set it (11 Indian languages and Indian English), or auto to detect it |
| End of turn | A turn-detection model, tuned with turn_detection | A fixed silence, endpointing_ms | Sarvam’s voice-activity detection: a pause set by turn_detection |
| Good for | Callers who mix Hindi and English; natural pauses | A known single language; strict control of the pause length | Indian languages beyond Hindi; Hindi agents that also speak with a Sarvam voice |
Sarvam writes Hindi in Devanagari, including English words said in Hindi (“अपॉइंटमेंट”). The LLM reads it fine, and transcripts show it as written.
Your workspace may not be allowed to use every provider. GET /api/v1/providers shows which providers your workspace can use; saving an agent with a provider it can’t use returns 422.
Turn detection
The agent starts its reply when the caller’s turn ends, so this setting trades speed against cutting the caller off.
turn_detection | Behaviour |
|---|---|
responsive | Answers quickly. Good for short answers (“yes”, “no”, a date). |
balanced | The default. |
patient | Waits longer. Good for slow speakers, elderly callers, or callers who think aloud. |
On Deepgram, raise endpointing_ms (for example to 500) if the agent talks over callers who pause mid-sentence, and lower it if replies feel slow.
If a caller pauses mid-sentence and the turn ends early, Vaakyo joins the two parts into one message for the LLM when the agent has not spoken in between. The transcript still shows them as two turns.
Fallback transcribers
transcriber.fallbacks lists up to two backup transcribers. When the transcriber fails to connect at the start of a call, or its connection drops during the call, the next one is connected and used for the rest of the call. Audio that arrives while it connects is kept and sent to it.
| Field | Type | Default | What it does |
|---|---|---|---|
provider | string | deepgram | cartesia, deepgram or sarvam. |
model | string | nova-3 | Replaced to fit the provider, as for the primary (ink-2, nova-3 or saaras:v3-realtime). |
language | string | hi | Deepgram and Sarvam, as for the primary. |
Backups use the primary’s keyterms, turn_detection and endpointing_ms. A backup can’t repeat the primary or an earlier backup (same provider, model and, on Deepgram and Sarvam, language); saving returns 422. Your workspace must be allowed to use the backup’s provider.
{
"transcriber": {
"provider": "cartesia",
"model": "ink-2",
"fallbacks": [{"provider": "deepgram", "model": "nova-3", "language": "multi"}]
}
}
Each switch is a fallback.used event on the call’s timeline, and the call’s providers.stt names the transcriber used last. Without fallbacks nothing changes: a transcriber that can’t connect fails the call, and one that stops mid-call ends it with an error. See Webhook events.
When the transcriber’s provider is at its concurrency limit as a call starts, the first fallback transcriber with room takes the call (a capacity.fallback event). See Webhook events.
Usage
The seconds of caller audio sent to the transcriber are counted in the call’s usage.stt_seconds.