Agents
Testing
Write test scenarios for an agent and run simulated text conversations against it before it goes live: an LLM plays the caller, your agent answers with its real prompt, model and tools, and an LLM judge scores the transcript against your criteria.
In the console, open an agent and pick the Tests tab. Over the API, scenarios and runs live under /api/v1/agents/{agent_id}/tests.
How a run works
-
The caller. An LLM role-plays the scenario’s
personalike a phone caller: short turns, in the language and mood the persona describes (Hindi or Hinglish too). -
The agent. Your agent’s saved version answers, text only, with the same system prompt a call gets (the scenario’s
variablesfilled in, the voice rules, its welcome message already said), its LLM provider, model and temperature, and its tools in a safe mode:Tool In a test Your HTTP tools Not called. The model gets the scenario’s stub from tool_stubs(default{"ok": true}) as{"status": 200, "result": <stub>}. Theirpre_call_messageis still said.search_knowledge_baseSearched for real (read-only). A knowledge base small enough is given whole in the prompt, as on a call. end_callEnds the conversation. reschedule_callRecorded, never booked. The conversation ends after the agent’s reply, as on a call. -
The end. The conversation stops when the agent hangs up, when the caller is done (after the agent’s last reply), or after
max_turnscaller turns. -
The judge. An LLM scores every criterion pass or fail with a one-line reason quoting the transcript, gives an overall
scorefrom 0 to 100, and writes a two or three sentencesummaryof what to change in the prompt.
A run passes when every required criterion passes. A scenario with no required criteria passes at a score of 70 or more.
Runs use the agent’s saved version (
agent_versionon the run). Save your changes before running tests.
The simulated caller is not a person, and the conversation is text, not audio: tests check what the agent says and does, not its voice, speech recognition or latency. Use a test call for those.
Scenarios
| Field | Type | Default | Meaning |
|---|---|---|---|
name | string | required | A short title, up to 120 characters. |
persona | string | required | Who the caller is, what they want, what details they have and how they behave. 10 to 4,000 characters. |
variables | object | {} | Values for the agent’s {variables}, as a call’s user_data gives them. |
max_turns | integer | 12 | The most caller turns, 1 to 30. |
criteria | array | [] | Up to 20 checks: {"text": "Asks for the caller's mohalla", "required": true}. A failed optional criterion only lowers the score. |
tool_stubs | object | {} | What each HTTP tool returns in tests, by tool name: {"create_ticket": {"ticket_id": "T-42"}}. |
An agent can have up to 50 scenarios. Deleting a scenario deletes its runs; deleting the agent deletes its scenarios.
curl -X POST https://api.vaakyo.com/api/v1/agents/$AGENT_ID/tests \
-H "X-API-Key: $VAAKYO_API_KEY" -H "Content-Type: application/json" \
-d '{
"name": "Angry caller from Shivaji Nagar",
"persona": "You are Ramesh, angry about garbage not collected for a week in Shivaji Nagar. You speak Hinglish and interrupt when the agent is slow.",
"variables": {"name": "Ramesh"},
"max_turns": 10,
"criteria": [
{"text": "Asks for the caller'\''s mohalla", "required": true},
{"text": "Never promises a date for the cleanup", "required": true},
{"text": "Ends politely", "required": false}
],
"tool_stubs": {"create_ticket": {"ticket_id": "T-42"}}
}'
GET /api/v1/agents/{agent_id}/tests lists the scenarios, each with last_run (id, status, score, cost_paise) of its newest run. GET, PUT and DELETE /api/v1/agents/{agent_id}/tests/{scenario_id} read, replace and delete one.
Suggested scenarios
Suggest scenarios (POST /api/v1/agents/{agent_id}/tests/suggest) asks an LLM for five scenarios with criteria, written from the agent’s saved prompt: the main job plus edge cases such as an unsure, impatient or off-topic caller. Nothing is saved: pick the ones you want and create them. The suggestion’s tokens are charged like a run.
{
"scenarios": [
{
"name": "Caller unsure of their mohalla",
"persona": "You are Sita, calling about a broken street light...",
"variables": {"name": "Sita"},
"max_turns": 8,
"criteria": [{"text": "Asks for a landmark when the caller is unsure", "required": true}],
"tool_stubs": {}
}
],
"model": "gemini-3.5-flash-lite",
"tokens": {"input": 1840, "output": 1210},
"cost_paise": 1
}
Running tests
# one scenario
curl -X POST https://api.vaakyo.com/api/v1/agents/$AGENT_ID/tests/$SCENARIO_ID/runs \
-H "X-API-Key: $VAAKYO_API_KEY"
# every scenario of the agent
curl -X POST https://api.vaakyo.com/api/v1/agents/$AGENT_ID/tests/run-all \
-H "X-API-Key: $VAAKYO_API_KEY"
Both answer 202 with runs in queued status (run-all returns {"batch_id", "runs"}). A workspace runs three tests at a time; the rest wait their turn. A run usually takes 30 seconds to two minutes and is stopped after five minutes. Poll GET /api/v1/test-runs/{run_id} until status is passed, failed or error. GET /api/v1/agents/{agent_id}/tests/{scenario_id}/runs lists a scenario’s runs, newest first, without transcripts; the last 50 are kept.
{
"id": "5f0c...",
"scenario_id": "a81e...",
"scenario_name": "Angry caller from Shivaji Nagar",
"agent_id": "c2d4...",
"agent_version": 7,
"status": "passed",
"turns": 4,
"ended_by": "agent",
"end_reason": "the agent ended the call",
"transcript": [
{"role": "agent", "text": "Namaste Ramesh, main Riya bol rahi hoon.", "turn": 0},
{"role": "caller", "text": "Mere mohalle mein hafte bhar se kachra pada hai!", "turn": 1},
{"role": "agent", "text": "Maaf kijiye. Aapka mohalla kaunsa hai?", "turn": 1},
{"role": "tool", "name": "create_ticket", "args": {"mohalla": "Shivaji Nagar"}, "result": {"status": 200, "result": {"ticket_id": "T-42"}}, "mode": "stubbed", "turn": 2}
],
"tool_calls": [{"name": "create_ticket", "args": {"mohalla": "Shivaji Nagar"}, "result": {"status": 200, "result": {"ticket_id": "T-42"}}, "mode": "stubbed", "turn": 2}],
"criteria": [
{"text": "Asks for the caller's mohalla", "required": true, "passed": true, "reason": "Asked \"Aapka mohalla kaunsa hai?\""},
{"text": "Never promises a date for the cleanup", "required": true, "passed": true, "reason": "Said \"date abhi nahi bata sakti\""},
{"text": "Ends politely", "required": false, "passed": false, "reason": "Hung up without a goodbye"}
],
"score": 82,
"summary": "Say a short goodbye before calling end_call. Acknowledge the caller's frustration once before asking for details.",
"usage": {
"agent": {"provider": "gemini", "model": "gemini-3.5-flash-lite", "input_tokens": 9800, "output_tokens": 190, "calls": 5, "cost_paise": 11},
"caller": {"provider": "gemini", "model": "gemini-3.5-flash-lite", "input_tokens": 2600, "output_tokens": 120, "calls": 4, "cost_paise": 4},
"judge": {"provider": "gemini", "model": "gemini-3.5-flash-lite", "input_tokens": 2900, "output_tokens": 310, "calls": 1, "cost_paise": 5}
},
"tokens": {"input": 15300, "output": 620},
"cost_paise": 20,
"error": "",
"created_at": "2026-10-04T15:30:00.120000+05:30",
"started_at": "2026-10-04T15:30:00.180000+05:30",
"finished_at": "2026-10-04T15:30:41.902000+05:30"
}
Tool entries have a mode: stubbed (an HTTP tool, not called), knowledge (a real knowledge base search), recorded (reschedule_call), end_call, or unknown (the model called a tool the agent doesn’t have). ended_by is agent, caller or max_turns. A run that could not finish has status: "error" and an error message; what it used so far is still charged.
Cost
Tests are charged from credits for the LLM tokens of all three roles at Vaakyo’s prices for their models: the agent at its own model’s price, the caller and the judge at the price of the models the Vaakyo team set for tests (Gemini 3.5 Flash-Lite by default). Each run is one test row in the credit ledger, and cost_paise on the run shows what it cost. A typical 12-turn run on Gemini 3.5 Flash-Lite costs well under ₹1 (about 40 paise); a larger agent model costs more. Runs and suggestions are refused with 402 when the workspace has no credits.
Permissions and Composer
Reading scenarios and runs needs agents.view; creating, changing, deleting, running and suggesting need agents.edit (runs cost credits).
Composer and MCP clients can work with tests through list_agent_tests, run_agent_test (one scenario, or all of them without scenario_id) and get_test_run. See MCP server.