From Plausible to Actionable: A Position on LLM Self-Explanations

· Source: cs.CL updates on arXiv.org · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Emerging Technologies & Innovation · Depth: Expert, extended

Summary

A position paper argues that Large Language Model (LLM) self-explanations, while often appearing plausible, are questionably faithful to the model's true reasoning process. Current XAI evaluation protocols for self-explanations have significant limitations, including their failure to consider stakeholder expertise or embrace explanation variation for plausibility, and their unsuitability for LLMs due to nondeterminism, prompt sensitivity, and reliance on internal knowledge for faithfulness. The paper advocates shifting the primary focus from plausibility and faithfulness to "actionability", defining it as the capacity of self-explanations to support informed decision-making and enable appropriate actions across diverse stakeholders. This actionable value is realized by positioning self-explanations as communicative interfaces for non-experts, tools for supporting human judgment, and facilitators of deliberation through diverse perspectives.

Key takeaway

For AI scientists and Machine Learning Engineers developing LLM-based systems for high-stakes applications, you should prioritize the "actionability" of self-explanations over their perceived faithfulness. Focus on designing explanations that serve as critical arguments for human decision-makers, communicate uncertainty, and facilitate diverse perspectives, rather than striving for unachievable faithful introspection. This approach mitigates automation bias and fosters informed trust, ensuring LLMs support, not replace, human judgment.

Key insights

LLM self-explanations are plausible but unfaithful; their true value lies in their actionability for informed decision-making.

Principles

Method

The paper proposes practical guidelines for assessing plausibility and faithfulness, then advocates a research agenda repositioning self-explanations as actionable communicative interfaces, decision-support tools, and deliberation facilitators.

In practice

Topics

Best for: Research Scientist, AI Scientist, Machine Learning Engineer, AI Ethicist

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by cs.CL updates on arXiv.org.