Auditing Alignment Controllability in LLMs via Political Axes

· Source: Computation and Language · Field: Technology & Digital — Artificial Intelligence & Machine Learning · Depth: Expert, quick

Summary

A study audited alignment controllability in seven leading large language models (LLMs), including GPT-5, Claude, Grok, Gemini, DeepSeek, Kimi, and Qwen. Researchers used a dispersion-first stress test. This involved 12 ideological personas, 70 Political Compass items, and ten replicates, generating 63,700 responses. The research found that contextual framing, primarily through system prompts, explains 88%-93% of variance on economic and society axes. Model identity accounts for under 3%. LLM responses are highly instruction-adjustable, though models shift differently and some saturate under extreme framings. The findings resolve prior conflicting audit results by highlighting non-centered baselines and the geometric nature of displacement. Political-coordinate audits therefore require steerability assessments reporting dispersion, symmetry, saturation, and refusal floors.

Key takeaway

For AI Ethicists and ML Engineers evaluating LLM alignment, recognize that models are highly steerable via prompts. Contextual framing dominates political response variance. Your audits must move beyond static political compass points to assess dispersion, symmetry, saturation, and refusal floors. This ensures a comprehensive understanding of a model's true behavioral range and potential for manipulation.

Key insights

LLM political alignment is highly steerable via prompts, with contextual framing explaining 88-93% of response variance.

Principles

Method

A dispersion-first stress test across 12 ideological personas, 70 Political Compass items, 10 replicates, and 7 leading LLMs (63,700 responses) to measure prompt-based controllability.

In practice

Topics

Best for: Research Scientist, AI Architect, AI Engineer, AI Scientist, Machine Learning Engineer, AI Ethicist

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Computation and Language.