WHBench: Evaluating Frontier LLMs with Expert-in-the-Loop Validation on Women's Health Topics
Summary
The Women's Health Benchmark (WHBench) introduces a new evaluation for large language models (LLMs) on women's health topics, addressing a critical gap in existing medical benchmarks. Comprising 47 expert-crafted clinical scenarios across 10 topics, WHBench targets specific LLM failure modes such as outdated guidelines, dosage errors, and missed health disparities. A panel of board-certified clinicians authored reference answers. The evaluation of 22 frontier, reasoning, and open-source models, using a 23-criterion rubric with asymmetric safety penalties, revealed significant performance limitations. Across 3,100 scored responses, no model achieved above 75% overall, with the top model, Claude Opus 4.6, scoring 72.1%. Crucially, even top models were fully correct in only 35.5% of cases, and harm rates ranged widely from 12.8% to 90.8%. All models demonstrated a universal weakness in integrating social determinants of health into clinical guidance, with pass rates between 0.7% and 19.1%.
Key takeaway
For machine learning engineers developing or deploying LLMs for clinical applications, particularly in women's health, you must recognize that even top models exhibit substantial safety risks and critical blind spots in health equity. Your current medical fine-tuning approaches may not address these nuanced failures. Prioritize integrating benchmarks like WHBench into your evaluation pipeline to rigorously test for specific failure modes, asymmetric safety penalties, and social determinants of health, ensuring your systems provide genuinely safe and equitable patient guidance.
Key insights
LLMs significantly underperform on women's health clinical scenarios, particularly in safety and health equity, despite strong general medical benchmark results.
Principles
- Clinical LLM evaluation requires failure-mode targeting and health equity criteria.
- High exam-style benchmark scores do not translate to open-ended clinical performance.
- LLM-as-judge can reliably rank systems, even with noisy individual response judgments.
Method
WHBench employs 47 expert-crafted, failure-mode-targeted clinical scenarios, scored by LLM judges using a 23-criterion rubric with asymmetric safety penalties and server-side score recalculation.
In practice
- Utilize WHBench to identify specific LLM capability gaps in women's health.
- Integrate safety and health equity criteria into clinical AI development pipelines.
- Mandate clinician review for all LLM-generated medical advice before patient use.
Topics
- Women's Health
- LLM Evaluation
- Clinical AI
- Health Equity
- Patient Safety
- Medical Benchmarking
Best for: Research Scientist, CTO, AI Architect, AI Scientist, Machine Learning Engineer, AI Ethicist
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by cs.CL updates on arXiv.org.