From Plausible to Actionable: A Position on LLM Self-Explanations
Summary
A position paper argues that Large Language Model (LLM) self-explanations, while often appearing plausible, are questionably faithful to the model's true reasoning process. Current XAI evaluation protocols for self-explanations have significant limitations, including their failure to consider stakeholder expertise or embrace explanation variation for plausibility, and their unsuitability for LLMs due to nondeterminism, prompt sensitivity, and reliance on internal knowledge for faithfulness. The paper advocates shifting the primary focus from plausibility and faithfulness to "actionability", defining it as the capacity of self-explanations to support informed decision-making and enable appropriate actions across diverse stakeholders. This actionable value is realized by positioning self-explanations as communicative interfaces for non-experts, tools for supporting human judgment, and facilitators of deliberation through diverse perspectives.
Key takeaway
For AI scientists and Machine Learning Engineers developing LLM-based systems for high-stakes applications, you should prioritize the "actionability" of self-explanations over their perceived faithfulness. Focus on designing explanations that serve as critical arguments for human decision-makers, communicate uncertainty, and facilitate diverse perspectives, rather than striving for unachievable faithful introspection. This approach mitigates automation bias and fosters informed trust, ensuring LLMs support, not replace, human judgment.
Key insights
LLM self-explanations are plausible but unfaithful; their true value lies in their actionability for informed decision-making.
Principles
- Plausibility evaluation must consider stakeholder expertise and embrace explanation variation.
- Faithfulness evaluation for LLMs must account for nondeterminism and prompt sensitivity.
- Self-explanations should be viewed as arguments, not literal accounts of model reasoning.
Method
The paper proposes practical guidelines for assessing plausibility and faithfulness, then advocates a research agenda repositioning self-explanations as actionable communicative interfaces, decision-support tools, and deliberation facilitators.
In practice
- Use self-explanations to translate complex XAI into natural language for non-experts.
- Design LLMs with safety protocols to generate arguments, not final decisions.
- Employ multiple LLMs as "advocates" to surface diverse viewpoints for deliberation.
Topics
- Large Language Models
- Explainable AI
- LLM Self-Explanations
- XAI Evaluation
- Actionability
- Human-AI Decision Support
Best for: Research Scientist, AI Scientist, Machine Learning Engineer, AI Ethicist
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by cs.CL updates on arXiv.org.