MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings

· Source: Computation and Language · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Robotics & Autonomous Systems · Depth: Expert, quick

Summary

MeetingToM is a new benchmark designed to evaluate Multimodal Large Language Models (MLLMs) on Theory-of-Mind (ToM) reasoning within naturalistic multi-party meetings. Current MLLMs face significant challenges in inferring others' beliefs and intentions, particularly when social cues are distributed across speech and behavior. Unlike existing benchmarks that focus on overt, verifiable signals, MeetingToM specifically targets complex social behaviors such as pseudo-consensus, where social pressure leads to apparent agreement despite private dissent. The benchmark is hierarchically structured to assess ToM at three levels: subject-level mental state prediction, dyadic-level addressee understanding, and group-level consensus reasoning. Systematic analyses using a unified evaluation protocol revealed persistent limitations in representative MLLMs regarding non-verbal cue integration, hidden attitude inference, and distinguishing genuine from pseudo-consensus, establishing MeetingToM as a critical testbed for advancing meeting-grounded ToM capabilities.

Key takeaway

For Machine Learning Engineers developing MLLMs for social interaction or meeting analysis, you should recognize current models' significant limitations in Theory-of-Mind reasoning. Your development efforts must prioritize integrating non-verbal cues and inferring hidden attitudes to distinguish genuine consensus from pseudo-consensus. This benchmark highlights the need for advanced models capable of understanding complex group dynamics, guiding your focus towards more robust and socially intelligent MLLM architectures.

Key insights

MeetingToM evaluates MLLMs' Theory-of-Mind in multi-party settings, revealing limitations in complex social reasoning.

Principles

Method

MeetingToM uses a hierarchical organization: subject-level mental state prediction, dyadic-level addressee understanding, and group-level consensus reasoning, with a unified evaluation protocol.

In practice

Topics

Best for: Research Scientist, AI Scientist, Machine Learning Engineer, NLP Engineer

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Computation and Language.