Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model

· Source: cs.CL updates on arXiv.org · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Emerging Technologies & Innovation, Scientific AI Interpretability · Depth: Expert, extended

Summary

The paper "Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model" by Markus J. Buehler investigates how the open-weight google/gemma-4-E4B-it model represents and uses materials science mechanisms. The study identifies three forms of information: concepts readable in individual hidden states, constitutive orientation in controlled state transformations, and causal control over engineering answers by selected internal representations. Using matched direct and Jacobian vocabulary readouts, a 60-law counterfactual benchmark, and causal interventions, the research found that Jacobian lenses reproduced concept ranks in 50 materials descriptions, identifying 9 of 10 mechanism families. While apparent physical organization was observed in hidden-state neighborhoods, a graph audit showed it was equally explained by numerical comparison. Crucially, state transformations correctly oriented 39 of 40 directional laws, and bidirectional interventions shifted answer probabilities towards physically appropriate outcomes across 12 matched cases. The study concludes that physical relationships are more visible in controlled state changes than in absolute states alone.

Key takeaway

For AI Scientists and ML Engineers evaluating LLM reliability in scientific domains, recognize that correct answers don't guarantee true physical understanding. Focus on analyzing internal state transformations, not just static representations, to verify constitutive law orientation. Implement causal interventions to confirm if internal directions predictably steer scientific decisions. This approach helps diagnose failures and build more robust, scientifically grounded AI systems.

Key insights

LLMs represent scientific mechanisms more reliably through state transformations than static states.

Principles

Method

Combines direct/Jacobian vocabulary readouts, option-free state geometry analysis, a 60-law counterfactual benchmark, and causal interventions to probe LLM internal states.

In practice

Topics

Code references

Best for: AI Scientist, Machine Learning Engineer, Research Scientist

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by cs.CL updates on arXiv.org.