D2VBench: Benchmarking Large Language Models with Value Dilemmas in Daily Scenarios
Summary
D2VBench is a new value alignment benchmark designed to evaluate Large Language Models (LLMs) in complex, real-world ethical dilemmas. It addresses limitations of prior benchmarks by offering 10,000 instances of daily scenarios, built through a multi-stage LLM-human collaboration and grounded in 158 fine-grained value concepts. The benchmark employs a hybrid evaluation paradigm, integrating multiple-choice and open-ended questions, which are scored based on five interpretability dimensions. Comprehensive evaluations on eight mainstream LLMs, including GPT-5.1, Gemini-3-pro-preview, and GLM-4.6, revealed D2VBench's high reliability and robustness. Results showed GPT-5.1 leading with an average score of approximately 65.6, but all LLMs struggled significantly with the Civilizational Progress category, and consistently demonstrated weakness in "Feasibility and Execution Difficulty".
Key takeaway
For AI Scientists and Machine Learning Engineers developing LLMs for real-world deployment, you should recognize that current models, even top performers like GPT-5.1, exhibit significant deficiencies in value alignment, particularly concerning abstract "Civilizational Progress" values and the "Feasibility and Execution Difficulty" of their proposed actions. Your development efforts must prioritize training LLMs to navigate complex, multi-value dilemmas and generate concrete, actionable solutions that consider real-world constraints, moving beyond principle-level reasoning.
Key insights
D2VBench offers a robust benchmark for LLM value alignment in complex, daily ethical dilemmas.
Principles
- Value alignment requires complex, multi-conflict daily scenarios.
- Hybrid evaluation captures LLM reasoning beyond surface responses.
- LLMs consistently underperform on practical execution feasibility.
Method
D2VBench constructs 10,000 dilemmas via LLM-human collaboration on 158 value concepts. It uses judge models to map open-ended LLM responses to options, scoring five interpretability dimensions with weighted universal value options.
In practice
- Prioritize LLM training on abstract "Civilizational Progress" values.
- Enhance LLM capabilities in "Feasibility and Execution Difficulty".
- Employ multi-label mapping for comprehensive response evaluation.
Topics
- Large Language Models
- LLM Value Alignment
- Ethical Benchmarking
- D2VBench Dataset
- Hybrid Evaluation
- Feasibility Analysis
Code references
Best for: Research Scientist, AI Scientist, Machine Learning Engineer, AI Ethicist
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by cs.CL updates on arXiv.org.