Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science
Summary
A study submitted on July 24, 2026, titled "Opaque Epistemic Mediation," investigated how commercial large language models (LLMs) evaluate ethnonationalist pseudo-science. Researchers tested four major LLM families—Claude, Grok, GPT, and Gemini—across four temporal snapshots from October 2025 to February 2026, using both API and web interfaces. Grok's Fast versions, which power the default X user experience, consistently assigned credibility scores of 70-75 to pseudo-scientific claims, significantly higher (two to five times) than other models, which scored 15-40. This pattern was absent for basic evolutionary consensus or refuted Lamarckian claims. The study also found a silent patch reversed Grok's behavior to stably high validation without documentation, and the same Grok model identifier yielded divergent API (75) and web (5.5) outputs. Furthermore, the defensible refusal to rate pseudo-science, observed in Claude Opus 4.1 (web) and GPT-5.1 Chat (API), eroded in successor versions. These findings suggest an LLM's epistemic stance is a contingent effect of opaque deployment configurations, not a stable model property.
Key takeaway
For AI Ethicists and Policy Makers evaluating LLM reliability, you must recognize that an LLM's stated "knowledge" is highly unstable and dependent on opaque deployment configurations, not just the underlying model. Your assessments should account for interface-specific behaviors, silent updates, and system prompt influences. Demand greater transparency from LLM providers regarding their deployment configurations and update practices to ensure epistemic accountability.
Key insights
LLM epistemic stances on contested science are unstable and opaque, driven by deployment configurations rather than inherent model properties.
Principles
- LLM epistemic stance is deployment-contingent.
- Silent updates can drastically alter LLM behavior.
- Interface routing affects LLM output credibility.
Method
Four major LLM families (Claude, Grok, GPT, Gemini) were tested on ethnonationalist pseudo-science via API and web interfaces across four temporal snapshots (Oct 2025-Feb 2026), with control prompts for comparison.
In practice
- Verify LLM outputs across interfaces.
- Monitor LLM behavior for undocumented changes.
- Scrutinize LLM credibility scores on contested topics.
Topics
- Large Language Models
- Epistemic Accountability
- Pseudo-science Validation
- Deployment Configurations
- Model Transparency
- AI Ethics
Best for: CTO, Research Scientist, VP of Engineering/Data, AI Scientist, AI Ethicist, Policy Maker
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by cs.CL updates on arXiv.org.