Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity
Summary
A study across three studies and six widely used Large Language Models (LLMs) investigated why LLM and human creativity evaluations converge or diverge. Findings indicate LLMs generally rely on a narrower subset of human creativity standards, showing strongest alignment in novelty assessment but clear divergence in contextual dimensions involving social, market, or reputational information. Each LLM exhibited distinct, model-specific standards varying in breadth. These differences impacted actual creativity judgments; LLM evaluations were moderately correlated with human ratings across 1,103 ideas, with broader-standard LLMs better distinguishing creative ideas. A third study with 1,195 ideas revealed LLMs were less sensitive to contextual information, which significantly altered human ratings but left LLM ratings largely unchanged.
Key takeaway
For Machine Learning Engineers deploying LLMs for creativity assessment, you must recognize that alignment with human judgment varies significantly. If your application requires evaluating ideas based on intrinsic qualities like novelty, LLMs may perform adequately. However, if contextual information (social, market, reputational) is crucial, your chosen LLM might be less sensitive, leading to divergent evaluations. Carefully select models based on the specific evaluation standards required for your task.
Key insights
LLM-human creativity evaluation alignment depends on judgment demands and the specific standards each model applies.
Principles
- LLMs use narrower creativity evaluation standards.
- Convergence on novelty, divergence on contextual data.
- Each LLM has distinct, model-specific standards.
Method
Three studies across six LLMs examined evaluation standards and their implications on creativity judgments, analyzing 1,103 and 1,195 ideas.
In practice
- Select LLM evaluators based on required judgment evidence.
- Consider LLM's sensitivity to contextual information.
- Recognize model-specific evaluation standards.
Topics
- Large Language Models
- Creativity Evaluation
- Human-AI Alignment
- Contextual Information
- Novelty Assessment
- Model Standards
Best for: Research Scientist, AI Scientist, Machine Learning Engineer, NLP Engineer
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by cs.CL updates on arXiv.org.