Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity
Summary
A study across three experiments and six large language models (LLMs) investigated the convergence and divergence of LLM creativity evaluations with human judgments. Findings indicate that LLMs generally rely on a narrower set of human creativity evaluation standards. Convergence with human standards was strongest in assessing novelty, while divergence was most evident in the contextual dimension, which incorporates social, market, and reputational information. Each LLM also exhibited distinct, model-specific standards varying in breadth. Study 2, involving 1,103 ideas, showed a moderate correlation between LLM and human evaluations, with LLMs possessing broader standards better distinguishing creative ideas. Study 3, with 1,195 ideas, revealed LLMs' reduced sensitivity to contextual information, which significantly impacted human ratings but left LLM ratings largely unchanged. This research explains mixed evidence on LLM-human alignment, suggesting it depends on the type of evidence required and the specific standards applied by each model.
Key takeaway
For machine learning engineers deploying LLMs for creativity assessment, recognize that your model choice significantly impacts evaluation outcomes. If your application requires nuanced judgments incorporating social, market, or reputational context, current LLMs may diverge from human perception. Prioritize models demonstrating broader evaluation standards for better alignment with human-judged creativity, especially when intrinsic qualities like novelty are paramount. Be aware that relying solely on LLMs for context-heavy creative tasks carries a risk of misaligned evaluations.
Key insights
LLM-human creativity evaluation alignment depends on model-specific standards and the contextual information demanded by the judgment.
Principles
- LLMs employ narrower creativity evaluation standards.
- Alignment with humans is strong for novelty, weak for context.
- Different LLMs apply distinct evaluation standards.
In practice
- LLMs are less sensitive to contextual information.
- Model choice impacts which ideas are deemed creative.
- Broader-standard LLMs better distinguish creative ideas.
Topics
- Large Language Models
- Creativity Evaluation
- Human-AI Alignment
- Contextual Information
- Novelty Assessment
- Model Selection
Best for: Research Scientist, AI Engineer, AI Product Manager, AI Scientist, Machine Learning Engineer, NLP Engineer
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Artificial Intelligence.