Beyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points

· Source: Artificial Intelligence · Field: Technology & Digital — Artificial Intelligence & Machine Learning · Depth: Expert, quick

Summary

The Epistemic Asymmetry Schelling Task (EAST) is introduced as a novel two-player dialogue game designed to benchmark robust and generalizable Theory of Mind (ToM) abilities in Large Language Models. This new evaluation method addresses limitations of traditional text-based ToM tests, such as the Sally-Anne task, which can be "gamed" due to LLM exposure during pre-training and may not reflect functional ToM in naturalistic settings. By requiring LLM-LLM dyads to independently converge on semantic Schelling points under varying states of epistemic transparency, EAST reveals a significant capability gap in functional social reasoning. Results indicate that only frontier models successfully navigate these tasks, with coordination failures primarily stemming from epistemic tracking errors, like conflating private and mutual knowledge. This highlights robust social reasoning and epistemic tracking as critical bottlenecks for LLM development, despite high performance on static benchmarks.

Key takeaway

For AI Scientists and Machine Learning Engineers focused on developing more human-like social reasoning in LLMs, you should recognize that traditional ToM benchmarks are insufficient. Your development efforts must prioritize improving epistemic tracking, specifically preventing models from conflating private and mutual knowledge. Integrating dialogue-based evaluation methods like EAST into your testing pipeline will provide a more robust assessment of functional social intelligence, guiding targeted improvements beyond superficial performance.

Key insights

The Epistemic Asymmetry Schelling Task (EAST) reveals LLMs' significant functional social reasoning and epistemic tracking gaps.

Principles

Method

EAST is a two-player dialogue game where LLM-LLM dyads independently converge on semantic Schelling points under varying epistemic transparency to evaluate robust Theory of Mind.

In practice

Topics

Best for: Research Scientist, AI Scientist, Machine Learning Engineer, NLP Engineer

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Artificial Intelligence.