A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents
Summary
A reliability assessment evaluates Gemini models as audio judges for scoring full-duplex agent conversations directly from raw stereo waveforms. The study tested Gemini 2.5 Flash, 3.5 Flash, and 3.1 Pro, using Gemini 2.5 Flash as the ground-truth model validated against three human raters on 209 stereo sessions. These sessions comprised 152 full-duplex conversations across 13 accent-and-condition strata and 57 adversarial defect-injected clips, scored on 8 production dimensions. Gemini 2.5 Flash demonstrated consistent reliability: its LALM-human Spearman rho departed from human-human rho by at most 0.07 on 5 of 8 dimensions, with 95 percent bootstrap confidence intervals overlapping on 7 of 8 dimensions. It agreed with the three-rater human mean within 1 point on 60 to 92 percent of sessions for 6 of 8 dimensions. Furthermore, the LALM was as sensitive as or better than humans in 45 of 48 (defect, dimension) cells. While 3.5 Flash improved simple agreement to 8 of 8 dimensions, 3.1 Pro rated several dimensions markedly lower. The authors estimate human rating costs two orders of magnitude more than the LALM workload.
Key takeaway
For MLOps Engineers evaluating full-duplex voice agents, consider deploying Gemini 2.5 Flash as an audio judge. This can drastically reduce evaluation costs by two orders of magnitude compared to human rating alone. You should specifically re-validate any model swap, like to Gemini 3.1 Pro, on calibration rather than assuming comparable performance from rank correlation alone, especially for critical dimensions.
Key insights
Gemini models can reliably serve as audio judges for full-duplex voice agent conversations, significantly reducing evaluation costs.
Principles
- LALM-human rank correlation can be comparable to human-human.
- Model swaps require re-validation, not just rank correlation.
Method
Gemini models score full-duplex agent conversations directly from raw stereo waveforms, validated against human raters on production dimensions and defect-injected clips.
In practice
- Deploy Gemini 2.5 Flash as a substitute or fourth rater for specific dimensions.
- Re-validate any model swap on calibration, not just rank correlation.
Topics
- Gemini models
- Full-duplex voice agents
- Audio judging
- Reliability assessment
- Model evaluation
- Speech analytics
Best for: Research Scientist, AI Scientist, MLOps Engineer, NLP Engineer
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Artificial Intelligence.