From Plausible to Actionable: A Position on LLM Self-Explanations

· Source: Computation and Language · Field: Technology & Digital — Artificial Intelligence & Machine Learning · Depth: Advanced, quick

Summary

Large Language Models (LLMs) can generate natural language self-explanations to rationalize their decisions, a promising direction for explainable artificial intelligence (XAI). While these explanations often appear plausible, their faithfulness to the model's true underlying reasoning remains uncertain. This analysis argues that self-explanations can be highly plausible, questionably faithful, yet profoundly actionable. It identifies limitations in standard evaluation protocols for LLM self-explanations. The paper also proposes practical guidelines for assessing plausibility and faithfulness. Furthermore, it advocates extending evaluation beyond these criteria to include actionability. This highlights how LLM rationalizations support informed decision-making for diverse stakeholders.

Key takeaway

For AI scientists and practitioners evaluating LLM outputs, recognize that LLM self-explanations, even if their faithfulness to internal reasoning is questionable, offer significant value for informed decision-making. You should prioritize evaluating these explanations based on their actionability and practical utility, rather than solely on their perceived faithfulness. This shift ensures that LLM rationalizations effectively support real-world applications and stakeholder needs.

Key insights

LLM self-explanations are valuable for decision-making and action, even with questionable faithfulness.

Principles

Method

Propose practical guidelines for assessing the plausibility and faithfulness of LLM-generated self-explanations.

In practice

Topics

Best for: Research Scientist, AI Product Manager, AI Scientist, AI Ethicist, Director of AI/ML

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Computation and Language.