D2VBench: Benchmarking Large Language Models with Value Dilemmas in Daily Scenarios

· Source: Computation and Language · Field: Technology & Digital — Artificial Intelligence & Machine Learning · Depth: Expert, quick

Summary

D2VBench is a new value alignment benchmark designed to evaluate large language models (LLMs) in real-world daily dilemma scenarios. It comprises 10,000 instances, constructed through a multi-stage collaboration between LLMs and humans, grounded in 158 manually annotated fine-grained value concepts. Addressing limitations of existing benchmarks, D2VBench employs a hybrid evaluation paradigm integrating multiple-choice and open-ended questions. Comprehensive evaluations on eight mainstream LLMs demonstrated its high reliability and robustness, effectively reflecting LLM alignment across various value categories and dimensions. This benchmark provides a more realistic and fine-grained tool for research into LLM value alignment, with the dataset publicly available.

Key takeaway

For AI Scientists or NLP Engineers developing or deploying LLMs, D2VBench offers a critical tool to assess value alignment beyond simplistic evaluations. You should integrate this benchmark into your model development lifecycle to identify and mitigate biases related to complex daily ethical dilemmas. Utilizing its hybrid evaluation paradigm will provide a more realistic and fine-grained understanding of your LLM's ethical performance, ensuring outputs are aligned with desired value concepts.

Key insights

D2VBench offers a robust benchmark for evaluating LLM value alignment in complex daily ethical dilemmas.

Principles

Method

D2VBench instances are constructed via multi-stage LLM-human collaboration, grounded in 158 manually annotated value concepts. Evaluation uses a hybrid multiple-choice and open-ended question paradigm.

In practice

Topics

Code references

Best for: Research Scientist, AI Scientist, NLP Engineer, AI Ethicist

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Computation and Language.