Agents
LLM tab
The LLM tab picks the model that decides what the agent says and calls its tools, and sets how varied and how long its replies are.
Fields
All fields live under llm in the agent object.
| Field | Type | Default | Allowed | What it does |
|---|---|---|---|---|
provider | string | gemini | gemini, bedrock, openai, azure_openai, anthropic, openrouter, deepseek | The LLM provider. Your workspace can use a provider only once the platform has its key and allows it for you (GET /api/v1/providers, usable). |
model | string | gemini-3.5-flash-lite | a model id from GET /api/v1/catalog/llm-models?provider=... | The model used for replies, the hang-up check and post-call analytics. For Azure OpenAI this is your deployment name. Changing provider without a model of that provider gives you the provider’s default model. |
temperature | number | 0.4 | 0.0 to 2.0 | Lower is more predictable; higher is more varied. |
max_tokens | integer | 300 | 16 to 4096 | The most tokens one reply can use. A reply that hits the limit is cut short. |
{
"llm": {
"provider": "gemini",
"model": "gemini-3.5-flash-lite",
"temperature": 0.4,
"max_tokens": 300
}
}
Providers
| Provider | provider | Default model | Models | Notes |
|---|---|---|---|---|
| Google Gemini | gemini | gemini-3.5-flash-lite | Gemini’s live list (text chat models) | The default. |
| AWS Bedrock | bedrock | us.amazon.nova-lite-v1:0 | Any model id, inference profile (us., eu., apac., global.) or ARN that supports tool use in the Converse API: Amazon Nova, Anthropic Claude, Meta Llama, Mistral | The profile’s region must match the platform key’s region, and the model must be enabled in that AWS account. The platform connects with a Bedrock API key or an IAM access key. Amazon Nova’s <thinking> notes are never spoken. |
| OpenAI | openai | gpt-4.1-mini | Known models plus OpenAI’s live chat models; any id | GPT-5 and o-series models run without a temperature, at minimal reasoning effort. |
| Azure OpenAI | azure_openai | gpt-4.1-mini | The platform’s configured deployments; any deployment name | model is the deployment name. |
| Anthropic | anthropic | claude-haiku-4-5 | Known models plus Anthropic’s live list; any id | Newer models (Opus/Sonnet 4.7 and later) take no temperature and run at low effort with room to think. |
| OpenRouter | openrouter | openai/gpt-4.1-mini | OpenRouter models whose endpoints support tools; any vendor/model id | Pick a model that supports tool calling: every agent has end_call. |
| DeepSeek | deepseek | deepseek-chat | deepseek-chat, deepseek-reasoner | Other ids are replaced by deepseek-chat. |
Every provider streams replies and calls tools the same way (see below). Gemini is still used by the platform itself for knowledge-base embeddings, the prompt writer and the Composer, whatever provider an agent uses.
Choosing a model
GET /api/v1/catalog/llm-models?provider=anthropic returns the models for one provider, and whether you may type any other id (custom_models):
{
"provider": "anthropic",
"models": [{"id": "claude-haiku-4-5", "label": "Claude Haiku 4.5"}, "..."],
"default_model": "claude-haiku-4-5",
"custom_models": true,
"model_label": "Model",
"model_hint": ""
}
GET /api/v1/catalog also lists the providers in llm_providers, and Gemini’s models in llm_models.
For Gemini, llm_models in GET /api/v1/catalog is the live list of text chat models (image, TTS, embedding and live models are left out).
curl https://api.vaakyo.com/api/v1/catalog -H "X-API-Key: $VAAKYO_API_KEY"
{
"llm_models": [{"id": "gemini-3.5-flash-lite", "label": "gemini-3.5-flash-lite"}, "..."],
"tts_models": [{"id": "sonic-3.6", "label": "Sonic 3.6 (latest)", "provider": "cartesia"}, "...", {"id": "bulbul:v3", "label": "Bulbul v3 (37 speakers, 11 languages)", "provider": "sarvam"}],
"stt_models": [{"id": "ink-2", "label": "Ink 2 (en, hi, fr, ja, es; detects language and turn end)", "provider": "cartesia"}, {"id": "nova-3", "label": "Nova 3", "provider": "deepgram"}, {"id": "saaras:v3-realtime", "label": "Saaras v3 realtime (Indian languages, live partial transcripts)", "provider": "sarvam"}, "..."],
"stt_providers": [{"id": "cartesia", "label": "Cartesia"}, {"id": "deepgram", "label": "Deepgram"}, {"id": "sarvam", "label": "Sarvam AI"}],
"deepgram_languages": [{"id": "hi", "label": "Hindi"}, {"id": "multi", "label": "Multilingual (switches between languages)"}, "..."],
"sarvam_stt_languages": [{"id": "auto", "label": "Detect automatically"}, {"id": "hi-IN", "label": "Hindi"}, "..."],
"sarvam_tts_languages": [{"id": "hi-IN", "label": "Hindi"}, {"id": "en-IN", "label": "English (India)"}, "..."],
"languages": [{"id": "hi", "label": "Hindi"}, {"id": "en", "label": "English"}]
}
On a phone call, the time from the caller finishing to the agent speaking matters more than anything else. Smaller “flash” models answer faster. Check first_audio_ms on your calls (see The call object) when you compare models.
How the model is used during a call
- Streaming. The reply streams in. Each finished sentence is sent to the voice at once, so the caller hears the first sentence while the model is still writing the rest.
- Tools. The model sees your tools, the built-in
end_call, andsearch_knowledge_basewhen the agent has knowledge bases (picked under Knowledge bases in this tab,knowledge_base_idsin the API). When it calls a tool, Vaakyo runs the HTTP request and gives the result back to the model, which then continues. A single caller turn allows up to four model requests, so the agent can chain a few tool calls. - History. The model sees the whole conversation so far. If the caller interrupted, it sees only the part of its reply the caller actually heard, followed by
…. - Errors. A rate limit, overload or timeout before the model has said anything is retried once. A model that refuses a parameter (for example
temperatureon newer reasoning models) is asked again without it. A model that hasn’t started its reply within 8 seconds is asked once more. If it still gives nothing, the agent says a short sorry (“Sorry, could you say that again?”, or a Hinglish version for Hindi voices), the call goes on, and anerrorevent is emitted. After 3 such failed replies in a row the call ends withhangup_by: "error"and statusfailed, and it is not charged. - Thinking. Gemini models think as little as the model allows (none on Gemini 2.5 Flash, “minimal” on Gemini 3 Flash models), so the first word comes sooner.
- Usage. Every request’s input and output tokens are added to the call’s
usage(llm_input_tokens,llm_output_tokens,llm_calls), including the hang-up check, and the call records the provider and model it used (providers.llm,providers.llm_model). The hang-up check and post-call analytics use the agent’s own provider and model; providers without a JSON mode get the expected JSON shape in the instruction. Post-call analytics is not counted inusage.
Tips
- Keep
max_tokenslow (200 to 400). Voice replies should be short, and the voice rules already ask for one to three sentences. - Keep
temperaturebetween0.2and0.6for agents that must follow a script or collect data.