Epistemic Stance Flexibility Probing: Measuring Prompt-Conditioned Register Shift in Large Language Models
Summary
The Epistemic Stance Flexibility Probing (ESFP) benchmark measures Large Language Models' (LLMs) ability to distinguish between externally attributed claims (what experts believe) and self-attributed claims (what the model believes), responding with appropriate epistemic registers. ESFP comprises 104 controlled items across six epistemic categories and five phrasing templates, evaluating responses on lexical self-attribution, representation-level responsiveness to role framing, sentence-level stance content density (assessed by an LLM judge panel), and cross-condition stance consistency. Evaluating eight frontier models from five vendors, the study found that epistemic flexibility is largely orthogonal to general model capability; a 27B open-weight model matched the strongest proprietary systems, and reasoning-optimized models did not consistently show higher flexibility. Stance content density provided the strongest signal, while surface-level lexical markers like "I think" often changed without corresponding shifts in expressed stance.
Key takeaway
For AI Scientists evaluating LLMs for conversational agents, recognize that general model capability does not guarantee coherent epistemic stance flexibility. You should integrate benchmarks like ESFP to specifically measure how models distinguish between externally attributed and self-attributed claims. This ensures your chosen models can reliably shift registers, enhancing trustworthiness beyond standard instruction following or accuracy metrics.
Key insights
LLM epistemic stance flexibility, distinguishing attributed vs. self-claims, is orthogonal to general model capability.
Principles
- Epistemic flexibility is largely orthogonal to general model capability.
- Surface-level lexical markers do not always reflect expressed stance.
Method
ESFP uses 104 items across six categories and five templates, evaluating lexical self-attribution, role framing responsiveness, LLM-judged stance content density, and cross-condition consistency.
In practice
- Use ESFP to assess LLM trustworthiness in conversational agents.
- Prioritize stance content density over lexical markers for evaluation.
Topics
- Epistemic Stance
- Large Language Models
- LLM Benchmarking
- Conversational AI
- Attribution
- Register Shift
Best for: Research Scientist, AI Engineer, AI Scientist, Machine Learning Engineer, NLP Engineer
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Computation and Language.