Response drift across frontier large language models

· Source: cs.CL updates on arXiv.org · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Data Science & Analytics · Depth: Expert, extended

Summary

A study involving 47 participants evaluated ten frontier large language models (LLMs) across 62 multi-domain questions, yielding 29,140 human assessments. It reveals that all LLMs exhibit "response drift," producing outputs deviating from expert-validated references. While two models, Claude (47.0% deviation) and Gemini (49.4% deviation), performed significantly better, eight others converged on a statistically indistinguishable ceiling (78–81% deviation). This drift is domain- and question-dependent, and automated similarity metrics explain less than 2% of human judgments, underscoring the need for human-centered evaluation. The evaluation period was December 2025 to February 2026.

Key takeaway

For AI Scientists and ML Engineers evaluating LLMs for knowledge-intensive applications, you must move beyond automated benchmarks. This study demonstrates that human-centered, reference-based fidelity evaluations are critical to accurately assess model performance and identify significant response drift. Relying solely on automated metrics will obscure substantial quality differences, leading to suboptimal model selection and deployment decisions.

Key insights

All frontier LLMs exhibit significant response drift, which automated metrics fail to detect, necessitating human evaluation.

Principles

Method

A fully crossed human evaluation with 47 participants assessing 10 LLMs on 62 multi-domain questions against expert references using a 5-point Likert scale.

In practice

Topics

Best for: AI Architect, AI Engineer, NLP Engineer, AI Scientist, Director of AI/ML, Machine Learning Engineer

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by cs.CL updates on arXiv.org.