Skip to content
VaakyoDocs
Navigation
Open console →

Agents

LLM tab

The LLM tab picks the model that decides what the agent says and calls its tools, and sets how varied and how long its replies are.

Fields

All fields live under llm in the agent object.

FieldTypeDefaultAllowedWhat it does
providerstringgeminigemini, bedrock, openai, azure_openai, anthropic, openrouter, deepseekThe LLM provider. Your workspace can use a provider only once the platform has its key and allows it for you (GET /api/v1/providers, usable).
modelstringgemini-3.5-flash-litea model id from GET /api/v1/catalog/llm-models?provider=...The model used for replies, the hang-up check and post-call analytics. For Azure OpenAI this is your deployment name. Changing provider without a model of that provider gives you the provider’s default model.
temperaturenumber0.40.0 to 2.0Lower is more predictable; higher is more varied.
max_tokensinteger30016 to 4096The most tokens one reply can use. A reply that hits the limit is cut short.
{
  "llm": {
    "provider": "gemini",
    "model": "gemini-3.5-flash-lite",
    "temperature": 0.4,
    "max_tokens": 300
  }
}

Providers

ProviderproviderDefault modelModelsNotes
Google Geminigeminigemini-3.5-flash-liteGemini’s live list (text chat models)The default.
AWS Bedrockbedrockus.amazon.nova-lite-v1:0Any model id, inference profile (us., eu., apac., global.) or ARN that supports tool use in the Converse API: Amazon Nova, Anthropic Claude, Meta Llama, MistralThe profile’s region must match the platform key’s region, and the model must be enabled in that AWS account. The platform connects with a Bedrock API key or an IAM access key. Amazon Nova’s <thinking> notes are never spoken.
OpenAIopenaigpt-4.1-miniKnown models plus OpenAI’s live chat models; any idGPT-5 and o-series models run without a temperature, at minimal reasoning effort.
Azure OpenAIazure_openaigpt-4.1-miniThe platform’s configured deployments; any deployment namemodel is the deployment name.
Anthropicanthropicclaude-haiku-4-5Known models plus Anthropic’s live list; any idNewer models (Opus/Sonnet 4.7 and later) take no temperature and run at low effort with room to think.
OpenRouteropenrouteropenai/gpt-4.1-miniOpenRouter models whose endpoints support tools; any vendor/model idPick a model that supports tool calling: every agent has end_call.
DeepSeekdeepseekdeepseek-chatdeepseek-chat, deepseek-reasonerOther ids are replaced by deepseek-chat.

Every provider streams replies and calls tools the same way (see below). Gemini is still used by the platform itself for knowledge-base embeddings, the prompt writer and the Composer, whatever provider an agent uses.

Choosing a model

GET /api/v1/catalog/llm-models?provider=anthropic returns the models for one provider, and whether you may type any other id (custom_models):

{
  "provider": "anthropic",
  "models": [{"id": "claude-haiku-4-5", "label": "Claude Haiku 4.5"}, "..."],
  "default_model": "claude-haiku-4-5",
  "custom_models": true,
  "model_label": "Model",
  "model_hint": ""
}

GET /api/v1/catalog also lists the providers in llm_providers, and Gemini’s models in llm_models.

For Gemini, llm_models in GET /api/v1/catalog is the live list of text chat models (image, TTS, embedding and live models are left out).

curl https://api.vaakyo.com/api/v1/catalog -H "X-API-Key: $VAAKYO_API_KEY"
{
  "llm_models": [{"id": "gemini-3.5-flash-lite", "label": "gemini-3.5-flash-lite"}, "..."],
  "tts_models": [{"id": "sonic-3.6", "label": "Sonic 3.6 (latest)", "provider": "cartesia"}, "...", {"id": "bulbul:v3", "label": "Bulbul v3 (37 speakers, 11 languages)", "provider": "sarvam"}],
  "stt_models": [{"id": "ink-2", "label": "Ink 2 (en, hi, fr, ja, es; detects language and turn end)", "provider": "cartesia"}, {"id": "nova-3", "label": "Nova 3", "provider": "deepgram"}, {"id": "saaras:v3-realtime", "label": "Saaras v3 realtime (Indian languages, live partial transcripts)", "provider": "sarvam"}, "..."],
  "stt_providers": [{"id": "cartesia", "label": "Cartesia"}, {"id": "deepgram", "label": "Deepgram"}, {"id": "sarvam", "label": "Sarvam AI"}],
  "deepgram_languages": [{"id": "hi", "label": "Hindi"}, {"id": "multi", "label": "Multilingual (switches between languages)"}, "..."],
  "sarvam_stt_languages": [{"id": "auto", "label": "Detect automatically"}, {"id": "hi-IN", "label": "Hindi"}, "..."],
  "sarvam_tts_languages": [{"id": "hi-IN", "label": "Hindi"}, {"id": "en-IN", "label": "English (India)"}, "..."],
  "languages": [{"id": "hi", "label": "Hindi"}, {"id": "en", "label": "English"}]
}

On a phone call, the time from the caller finishing to the agent speaking matters more than anything else. Smaller “flash” models answer faster. Check first_audio_ms on your calls (see The call object) when you compare models.

How the model is used during a call

  • Streaming. The reply streams in. Each finished sentence is sent to the voice at once, so the caller hears the first sentence while the model is still writing the rest.
  • Tools. The model sees your tools, the built-in end_call, and search_knowledge_base when the agent has knowledge bases (picked under Knowledge bases in this tab, knowledge_base_ids in the API). When it calls a tool, Vaakyo runs the HTTP request and gives the result back to the model, which then continues. A single caller turn allows up to four model requests, so the agent can chain a few tool calls.
  • History. The model sees the whole conversation so far. If the caller interrupted, it sees only the part of its reply the caller actually heard, followed by ….
  • Errors. A rate limit, overload or timeout before the model has said anything is retried once. A model that refuses a parameter (for example temperature on newer reasoning models) is asked again without it. A model that hasn’t started its reply within 8 seconds is asked once more. If it still gives nothing, the agent says a short sorry (“Sorry, could you say that again?”, or a Hinglish version for Hindi voices), the call goes on, and an error event is emitted. After 3 such failed replies in a row the call ends with hangup_by: "error" and status failed, and it is not charged.
  • Thinking. Gemini models think as little as the model allows (none on Gemini 2.5 Flash, “minimal” on Gemini 3 Flash models), so the first word comes sooner.
  • Usage. Every request’s input and output tokens are added to the call’s usage (llm_input_tokens, llm_output_tokens, llm_calls), including the hang-up check, and the call records the provider and model it used (providers.llm, providers.llm_model). The hang-up check and post-call analytics use the agent’s own provider and model; providers without a JSON mode get the expected JSON shape in the instruction. Post-call analytics is not counted in usage.

Tips

  • Keep max_tokens low (200 to 400). Voice replies should be short, and the voice rules already ask for one to three sentences.
  • Keep temperature between 0.2 and 0.6 for agents that must follow a script or collect data.
Esc