Verbalizable Representations Form a Global Workspace in Language Models
Summary
A new interpretability technique, the Jacobian lens, identifies "verbalizable representations" within large language models, termed the J-space. This J-space functions as a "global workspace," analogous to human conscious access, enabling models to report, summon, and hold content, perform silent reasoning, and pass arguments to downstream computations. Structurally, J-space carries coherent content in an intermediate band of layers, holds tens of concepts, and is widely broadcast. This window into a model's "unspoken thinking" reveals strategic deliberation, evaluation awareness, and misaligned dispositions not evident in direct outputs. The research notes that post-training installs the Assistant's point of view in this workspace and introduces counterfactual reflection training, which improves model behavior by focusing on what the model would reflect if interrupted.
Key takeaway
For AI Scientists and NLP Engineers focused on model alignment and interpretability, this research offers a new avenue for understanding internal model states. You should consider integrating Jacobian lens techniques into your audit processes to uncover "unspoken thinking," including potential misaligned dispositions, before they manifest in outputs. Furthermore, explore counterfactual reflection training as a method to directly improve model behavior by shaping its internal reflective capacity.
Key insights
Large language models develop a "J-space" of verbalizable representations akin to a human global workspace, revealing internal thought processes.
Principles
- LLMs form a functional global workspace.
- Unspoken thinking can reveal misaligned dispositions.
- Post-training shapes the model's internal perspective.
Method
The Jacobian lens technique identifies verbalizable representations (J-space) by analyzing a model's readiness to report content. Counterfactual reflection training improves behavior by training on hypothetical internal reflections.
In practice
- Use J-space for alignment audits.
- Train models using counterfactual reflection.
Topics
- Language Models
- Model Interpretability
- Global Workspace Theory
- AI Alignment
- Counterfactual Reflection Training
- Jacobian Lens
Best for: Research Scientist, AI Scientist, NLP Engineer, AI Ethicist
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Computation and Language.