Response drift across frontier large language models
Summary
A study involving 47 participants evaluated ten frontier large language models (LLMs) across 62 multi-domain questions, yielding 29,140 human assessments. It reveals that all LLMs exhibit "response drift," producing outputs deviating from expert-validated references. While two models, Claude (47.0% deviation) and Gemini (49.4% deviation), performed significantly better, eight others converged on a statistically indistinguishable ceiling (78–81% deviation). This drift is domain- and question-dependent, and automated similarity metrics explain less than 2% of human judgments, underscoring the need for human-centered evaluation. The evaluation period was December 2025 to February 2026.
Key takeaway
For AI Scientists and ML Engineers evaluating LLMs for knowledge-intensive applications, you must move beyond automated benchmarks. This study demonstrates that human-centered, reference-based fidelity evaluations are critical to accurately assess model performance and identify significant response drift. Relying solely on automated metrics will obscure substantial quality differences, leading to suboptimal model selection and deployment decisions.
Key insights
All frontier LLMs exhibit significant response drift, which automated metrics fail to detect, necessitating human evaluation.
Principles
- Response drift is universal across LLMs.
- Automated metrics do not capture human-perceived fidelity.
- Model choice primarily determines drift magnitude.
Method
A fully crossed human evaluation with 47 participants assessing 10 LLMs on 62 multi-domain questions against expert references using a 5-point Likert scale.
In practice
- Implement reference-based fidelity assessments.
- Consider domain-specific LLM routing.
- Avoid sole reliance on automated NLP metrics.
Topics
- Large Language Models
- Human Evaluation
- Response Drift
- Model Fidelity
- Automated Benchmarks
- Evaluation Metrics
Best for: AI Architect, AI Engineer, NLP Engineer, AI Scientist, Director of AI/ML, Machine Learning Engineer
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by cs.CL updates on arXiv.org.