Posts by Category

Critical Analysis

General reading group

Paper notes: A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models (Rai et al., 2025)

1 minute read

Published:

On Friday the 5th of September, the general reading group continued its mechanistic interpretability sprint with A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models. The research team comes from across several US universities, with one member from Salesforce Research. Its lead author is a PhD student, and its second author a PhD working in industry. The survey was posted on arXiv and presented as a tutorial at ICML 2025.

Paper notes: On the Biology of a Large Language Model (Lindsey et al., 2025)

2 minute read

Published:

Last week, Deep Network’s reading group read On the Biology of a Large Language Model. The research team comes from Anthropic’s interpretability research group, and was published in the Transformer Circuits interactive research thread as well as on the Anthropic website.

Mathematics reading group

Paper Notes: Emergent Cooperation and Strategy Adaptation in Multi-Agent Systems: An Extended Coevolutionary Theory with LLMs (Zarzà et al., 2023)

4 minute read

Published:

Welcome to Paper Notes, where we record our groups’ weekly discussions of innovative papers from across artificial intelligence. On Tuesday the 21st of October, our mathematics reading group read Emergent Cooperation and Strategy Adaptation in Multi-Agent Systems: An Extended Coevolutionary Theory with LLMs, published in MDPI Electronics in 2023. The authors come from a collection of three universities and a lab, across Spain and Germany.

Paper Notes

Paper Notes: Emergent Cooperation and Strategy Adaptation in Multi-Agent Systems: An Extended Coevolutionary Theory with LLMs (Zarzà et al., 2023)

4 minute read

Published:

Welcome to Paper Notes, where we record our groups’ weekly discussions of innovative papers from across artificial intelligence. On Tuesday the 21st of October, our mathematics reading group read Emergent Cooperation and Strategy Adaptation in Multi-Agent Systems: An Extended Coevolutionary Theory with LLMs, published in MDPI Electronics in 2023. The authors come from a collection of three universities and a lab, across Spain and Germany.

Paper notes: A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models (Rai et al., 2025)

1 minute read

Published:

On Friday the 5th of September, the general reading group continued its mechanistic interpretability sprint with A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models. The research team comes from across several US universities, with one member from Salesforce Research. Its lead author is a PhD student, and its second author a PhD working in industry. The survey was posted on arXiv and presented as a tutorial at ICML 2025.

Paper notes: On the Biology of a Large Language Model (Lindsey et al., 2025)

2 minute read

Published:

Last week, Deep Network’s reading group read On the Biology of a Large Language Model. The research team comes from Anthropic’s interpretability research group, and was published in the Transformer Circuits interactive research thread as well as on the Anthropic website.