A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents

· Source: Artificial Intelligence · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Robotics & Autonomous Systems · Depth: Expert, quick

Summary

A reliability assessment evaluates Gemini models as audio judges for scoring full-duplex agent conversations directly from raw stereo waveforms. The study tested Gemini 2.5 Flash, 3.5 Flash, and 3.1 Pro, using Gemini 2.5 Flash as the ground-truth model validated against three human raters on 209 stereo sessions. These sessions comprised 152 full-duplex conversations across 13 accent-and-condition strata and 57 adversarial defect-injected clips, scored on 8 production dimensions. Gemini 2.5 Flash demonstrated consistent reliability: its LALM-human Spearman rho departed from human-human rho by at most 0.07 on 5 of 8 dimensions, with 95 percent bootstrap confidence intervals overlapping on 7 of 8 dimensions. It agreed with the three-rater human mean within 1 point on 60 to 92 percent of sessions for 6 of 8 dimensions. Furthermore, the LALM was as sensitive as or better than humans in 45 of 48 (defect, dimension) cells. While 3.5 Flash improved simple agreement to 8 of 8 dimensions, 3.1 Pro rated several dimensions markedly lower. The authors estimate human rating costs two orders of magnitude more than the LALM workload.

Key takeaway

For MLOps Engineers evaluating full-duplex voice agents, consider deploying Gemini 2.5 Flash as an audio judge. This can drastically reduce evaluation costs by two orders of magnitude compared to human rating alone. You should specifically re-validate any model swap, like to Gemini 3.1 Pro, on calibration rather than assuming comparable performance from rank correlation alone, especially for critical dimensions.

Key insights

Gemini models can reliably serve as audio judges for full-duplex voice agent conversations, significantly reducing evaluation costs.

Principles

Method

Gemini models score full-duplex agent conversations directly from raw stereo waveforms, validated against human raters on production dimensions and defect-injected clips.

In practice

Topics

Best for: Research Scientist, AI Scientist, MLOps Engineer, NLP Engineer

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Artificial Intelligence.