Psychological Competence as a Missing Dimension in AI Evaluation

· Source: cs.AI updates on arXiv.org · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Human-AI Interaction · Depth: Advanced, quick

Summary

A new dimension, "psychological competence," is proposed for evaluating human-facing AI systems, addressing a gap in current frameworks that prioritize technical performance like accuracy and robustness. This concept defines an AI system's capacity to appropriately support user cognition, emotional interpretation, and behavioral decision-making, considering the user, context, and interaction purpose. As AI increasingly serves as advisors, coaches, and companions, its influence on user reasoning, beliefs, trust, and decisions necessitates evaluating the human-AI interaction, not just the model itself. Key interaction properties include framing, tone, perceived authority, responsiveness, uncertainty handling, and conversational guidance. The paper outlines a conceptual framework and suggests assessment through scenario-based probes, structured human evaluation, and model-assisted methods, advocating for its adoption by model providers, organizations, researchers, and regulators.

Key takeaway

For research scientists and model providers developing human-facing AI, recognize that traditional technical performance metrics are insufficient. Your evaluation frameworks must integrate psychological competence, assessing how AI influences user cognition, emotions, and decision-making. Implement scenario-based probes and structured human evaluations to directly measure interaction properties like framing and perceived authority, ensuring your systems support users appropriately and ethically.

Key insights

AI evaluation must expand beyond technical performance to include psychological competence, assessing human-AI interaction effects on users.

Principles

Method

Psychological competence may be assessed via scenario-based probes, structured human evaluation, and model-assisted evaluation methods. This clarifies the construct's boundaries and conceptual framework.

In practice

Topics

Best for: AI Product Manager, AI Scientist, Research Scientist, AI Ethicist

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by cs.AI updates on arXiv.org.