Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity

· Source: Artificial Intelligence · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Data Science & Analytics · Depth: Expert, quick

Summary

A study across three experiments and six large language models (LLMs) investigated the convergence and divergence of LLM creativity evaluations with human judgments. Findings indicate that LLMs generally rely on a narrower set of human creativity evaluation standards. Convergence with human standards was strongest in assessing novelty, while divergence was most evident in the contextual dimension, which incorporates social, market, and reputational information. Each LLM also exhibited distinct, model-specific standards varying in breadth. Study 2, involving 1,103 ideas, showed a moderate correlation between LLM and human evaluations, with LLMs possessing broader standards better distinguishing creative ideas. Study 3, with 1,195 ideas, revealed LLMs' reduced sensitivity to contextual information, which significantly impacted human ratings but left LLM ratings largely unchanged. This research explains mixed evidence on LLM-human alignment, suggesting it depends on the type of evidence required and the specific standards applied by each model.

Key takeaway

For machine learning engineers deploying LLMs for creativity assessment, recognize that your model choice significantly impacts evaluation outcomes. If your application requires nuanced judgments incorporating social, market, or reputational context, current LLMs may diverge from human perception. Prioritize models demonstrating broader evaluation standards for better alignment with human-judged creativity, especially when intrinsic qualities like novelty are paramount. Be aware that relying solely on LLMs for context-heavy creative tasks carries a risk of misaligned evaluations.

Key insights

LLM-human creativity evaluation alignment depends on model-specific standards and the contextual information demanded by the judgment.

Principles

In practice

Topics

Best for: Research Scientist, AI Engineer, AI Product Manager, AI Scientist, Machine Learning Engineer, NLP Engineer

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Artificial Intelligence.