WHBench: Evaluating Frontier LLMs with Expert-in-the-Loop Validation on Women's Health Topics

· Source: cs.CL updates on arXiv.org · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Data Science & Analytics, Medical AI Evaluation · Depth: Expert, extended

Summary

The Women's Health Benchmark (WHBench) introduces a new evaluation for large language models (LLMs) on women's health topics, addressing a critical gap in existing medical benchmarks. Comprising 47 expert-crafted clinical scenarios across 10 topics, WHBench targets specific LLM failure modes such as outdated guidelines, dosage errors, and missed health disparities. A panel of board-certified clinicians authored reference answers. The evaluation of 22 frontier, reasoning, and open-source models, using a 23-criterion rubric with asymmetric safety penalties, revealed significant performance limitations. Across 3,100 scored responses, no model achieved above 75% overall, with the top model, Claude Opus 4.6, scoring 72.1%. Crucially, even top models were fully correct in only 35.5% of cases, and harm rates ranged widely from 12.8% to 90.8%. All models demonstrated a universal weakness in integrating social determinants of health into clinical guidance, with pass rates between 0.7% and 19.1%.

Key takeaway

For machine learning engineers developing or deploying LLMs for clinical applications, particularly in women's health, you must recognize that even top models exhibit substantial safety risks and critical blind spots in health equity. Your current medical fine-tuning approaches may not address these nuanced failures. Prioritize integrating benchmarks like WHBench into your evaluation pipeline to rigorously test for specific failure modes, asymmetric safety penalties, and social determinants of health, ensuring your systems provide genuinely safe and equitable patient guidance.

Key insights

LLMs significantly underperform on women's health clinical scenarios, particularly in safety and health equity, despite strong general medical benchmark results.

Principles

Method

WHBench employs 47 expert-crafted, failure-mode-targeted clinical scenarios, scored by LLM judges using a 23-criterion rubric with asymmetric safety penalties and server-side score recalculation.

In practice

Topics

Best for: Research Scientist, CTO, AI Architect, AI Scientist, Machine Learning Engineer, AI Ethicist

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by cs.CL updates on arXiv.org.