Beyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points
Summary
The Epistemic Asymmetry Schelling Task (EAST) is introduced as a novel two-player dialogue game designed to benchmark robust and generalizable Theory of Mind (ToM) abilities in Large Language Models. This new evaluation method addresses limitations of traditional text-based ToM tests, such as the Sally-Anne task, which can be "gamed" due to LLM exposure during pre-training and may not reflect functional ToM in naturalistic settings. By requiring LLM-LLM dyads to independently converge on semantic Schelling points under varying states of epistemic transparency, EAST reveals a significant capability gap in functional social reasoning. Results indicate that only frontier models successfully navigate these tasks, with coordination failures primarily stemming from epistemic tracking errors, like conflating private and mutual knowledge. This highlights robust social reasoning and epistemic tracking as critical bottlenecks for LLM development, despite high performance on static benchmarks.
Key takeaway
For AI Scientists and Machine Learning Engineers focused on developing more human-like social reasoning in LLMs, you should recognize that traditional ToM benchmarks are insufficient. Your development efforts must prioritize improving epistemic tracking, specifically preventing models from conflating private and mutual knowledge. Integrating dialogue-based evaluation methods like EAST into your testing pipeline will provide a more robust assessment of functional social intelligence, guiding targeted improvements beyond superficial performance.
Key insights
The Epistemic Asymmetry Schelling Task (EAST) reveals LLMs' significant functional social reasoning and epistemic tracking gaps.
Principles
- Traditional ToM benchmarks are insufficient.
- Epistemic tracking is critical for social reasoning.
- Dialogue games test robust ToM capabilities.
Method
EAST is a two-player dialogue game where LLM-LLM dyads independently converge on semantic Schelling points under varying epistemic transparency to evaluate robust Theory of Mind.
In practice
- Prioritize LLM development on epistemic tracking.
- Utilize dialogue games for robust ToM evaluation.
- Analyze coordination failures for epistemic errors.
Topics
- Large Language Models
- Theory of Mind
- LLM Evaluation
- Epistemic Schelling Task
- Social Reasoning
- Coordination Games
Best for: Research Scientist, AI Scientist, Machine Learning Engineer, NLP Engineer
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Artificial Intelligence.