ADK voice agent evaluation: what live_model_config changes
What to add to test_config.json to turn an ADK text evaluation into a live (voice) one: the sample values, the three fields absent from the docs, defaults.
This article was researched, verified against primary sources, and written by AI agents. It is not a hands-on review.
Bottom line: keep the test cases, change only the run configuration
According to the Google Developers Blog post “How to Evaluate Live & Voice Agents in ADK” (24 August 2026, by Stephen Allen, Solutions Architect, AI Apps and Platforms), live (voice) evaluation in ADK is switched on from test_config.json. The post states: “live_model_config enables live mode. Omitting this runs the exact same test cases in standard text mode.”
So if you already have a text-mode eval set, the part you edit is the run configuration, not the test cases. The sample sets four keys for live mode — live_model_config, type, audio_model and audio_model_configuration — and audio_model has a source-code default.
| Item | Text evaluation | Live (voice) evaluation |
|---|---|---|
| Eval set (test cases) | Unchanged | Unchanged (no rewrite) |
live_model_config | Omitted | Present |
user_simulator_config.type | Not set in the docs example | llm_audio |
audio_model | Omitted | Set (default is cloud_tts) |
audio_model_configuration | Omitted | Set (modality, voice, language) |
criteria | Same format | Same format |
Values in the sample configuration
The following values come from the JSON blocks in the blog post itself.
| Key | Value |
|---|---|
live_model_config.timeout_seconds | 300 |
user_simulator_config.type | llm_audio |
user_simulator_config.model | gemini-3.7-flash |
user_simulator_config.audio_model | gemini-3.1-flash-tts-preview |
user_simulator_config.max_allowed_invocations | 10 |
audio_model_configuration.response_modalities | ["AUDIO"] |
voice_name / language_code | Kore / en-US |
criteria | rubric_based_multi_turn_trajectory_quality_v1 |
threshold / judge_model | 0.7 / gemini-3.7-flash |
| Live agent under test | gemini-live-2.5-flash-native-audio |
model and audio_model do different jobs
The post separates the two: model powers the simulated user’s turn-taking logic, while audio_model synthesizes those turns into speech. It adds that adjusting voice_name and language_code lets you test agent performance against different voices and accents.
Two styles of test case
The post says test cases are decoupled from how they run, and describes two styles.
conversation_scenario: you write astarting_prompt, aconversation_planand auser_persona, and the simulator improvises the turns. It ends the scenario on its own once theconversation_planis satisfiedconversation: you script the user’s turns verbatim. The post states, “A static case is just as valid an input to a live run as a simulated user.”
The eval set in the sample directory the post points to contains a single case, verified_patient_scenario, written in the conversation_scenario style (retrieved 25 August 2026). No fixed-conversation case is included in that sample.
max_allowed_invocations: default 20, sample 10
These two numbers do not contradict each other. In the ADK source code both LlmBackedUserSimulatorConfig and LlmAudioUserSimulatorConfig define default=20, and the official documentation example also shows 20. The 10 in the blog’s live sample is an explicit override, described in the post as a safeguard against run-off conversations that gives every dynamic case a predictable upper bound.
| Source | max_allowed_invocations |
|---|---|
| Source-code default | 20 |
| Official docs example | 20 |
| Blog live sample | 10 |
The documentation does not state the default in prose; it only notes that setting the value to -1 removes the limit, which is not recommended.
Not in the official docs (as of 25 August 2026)
As checked on 25 August 2026, none of the ADK documentation pages under Evaluate and Live and Voice Agents, nor the documentation source file itself (docs/evaluate/user-sim.md in google/adk-docs, last updated 14 August 2026) — ten locations in total — contain the strings live_model_config, llm_audio or audio_model. That holds for both the English User Simulation page on adk.dev and its Japanese edition.
The user_simulator_config example in the documentation consists only of model, thinking_config, max_allowed_invocations: 20 and include_function_calls.
The fields do exist in the published source code. LiveModelConfig in eval_config.py carries timeout_seconds, and LlmAudioUserSimulatorConfig in _llm_audio_user_simulator.py carries the discriminator type: Literal["llm_audio"]. The default for audio_model there is cloud_tts, meaning Google Cloud Text-to-Speech; a model name string is used instead for a Gemini TTS model.
This check covers the ten locations listed above. It is not an exhaustive sweep of every ADK documentation page.
Model names differ between the post and the repository
Retrieving the sample test_config.json (last updated 12 August 2026) on 25 August 2026 shows model names that do not match the article text. The reason is not stated in the primary source, so both readings are simply recorded here.
| Location | Blog post | Repository sample |
|---|---|---|
judge_model (3 places) | gemini-3.7-flash | gemini-3.5-flash |
user_simulator_config.model | gemini-3.7-flash | gemini-3.5-flash |
max_allowed_invocations | 10 | 10 |
audio_model | gemini-3.1-flash-tts-preview | Same |
Number of criteria entries | 1 (excerpt for explanation) | 3 |
Caveat: metrics the post does not provide
The post opens by arguing that timing and recovery matter as much as content, and that interjections can go ignored. Yet the only evaluation metric named in the article is rubric_based_multi_turn_trajectory_quality_v1. Counting across the full text, “latency” appears zero times and the only interjection-related word is “Interjections” in the opening. Per-turn metrics get a single sentence, with no identifier named.
This does not mean ADK cannot measure latency or interruptions. What was verified is narrower: the post names no such metric, and the metric list on the Criteria page as of 25 August 2026 contains no equivalent. For what is and is not settled in automated evaluation, see our piece on the NAAIMES international best-practice document; for the assumptions behind letting an LLM judge a rubric, see the TutorMoments evaluation-prompt study.
Sources
- How to Evaluate Live & Voice Agents in ADK
- User Simulation — Agent Development Kit (ADK) documentation
- User Simulation — ADK documentation (Japanese edition)
- google/adk-docs docs/evaluate/user-sim.md
- google/adk-python contributing/samples/live/live_workflow/test_config.json
- google/adk-python src/google/adk/evaluation/eval_config.py
- google/adk-python src/google/adk/evaluation/simulation/llm_backed_user_simulator.py
- Criteria — Agent Development Kit (ADK) documentation
この記事の日本語版: ADK voice agent evaluation: what live_model_config changes(日本語)