Paper notes: A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models (Rai et al., 2025)
Published:
On Friday the 5th of September, the general reading group continued its mechanistic interpretability sprint with A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models.1 The research team comes from across several US universities, with one member from Salesforce Research. Its lead author is a PhD student, and its second author a PhD working in industry. The survey was posted on arXiv and presented as a tutorial at ICML 2025.2
The preprint surveys mechanistic interpretability, with a focus on the workflow of MI research and open problems in the field.
Hypotheses on the paper
This week we tried something a little different. Rather than discussing the paper, we set ourselves the challenge of generating hypotheses about its contents. Here are some of these hypotheses:
- Observation: most MI techniques exclude feedforward, layernorm, and embedding matrices from their analysis.
- Hypothesis: a hybrid method analysing feedforward, layernorm, and embedding matrices would provide a more faithful representation of causal mechanisms in the model.
- Counter: this might require relaxing simplifying assumptions about representations (eg linear representation hypothesis).
- Observation: superposition is prevalent in LLMs.
- Hypothesis: superposition hurts robustness of behaviours targeted during post-training.
- Observation: circular representations are more information-efficient for encoding cyclical relationships, like time.
- Hypothesis: models learn trigonometric operations to process cyclical information.
- Observation: the SHIFT technique removes non-human-interpretable features to improve generalisation.
- Hypothesis: human-AI collaboration on a task would be easier when using SHIFT.
- Hypothesis: models confidence measures would be more accurate when using SHIFT.
- Observation: the model encodes the truth of context tokens in a feature.
- Hypothesis: we can detect lying despite hidden misalignment using truth features.
- Hypothesis: LLMs have an internal “truth uncertainty” feature more accurate, reliable, or robust to misalignment than logits.
- Hypothesis: we could model misaligned behaviour from MI features to detect misaligned behaviour at runtime.
This blog was originally posted at https://deep-network.org/2025/09/23/paper-notes-a-practical-review-of-mechanistic-interpretability-for-transformer-based-language-models-rai-et-al-2025/
Rai, D., Zhou, Y., Feng, S., Saparov, A., & Yao, Z. (2024). A practical review of mechanistic interpretability for transformer-based language models. arXiv preprint arXiv:2407.02646. https://arxiv.org/abs/2407.02646 ↩
Rai, D., Zhou, Y., Feng, S., Saparov, A., & Yao, Z. (2025, July 13). ICML 2025 Tutorial on Mechanistic Interpretability for Language Models. https://ziyu-yao-nlp-lab.github.io/ICML25-MI-Tutorial.github.io/ ↩
