An Introduction to Mechanistic Interpretability – Neel Nanda IASEAI 2025
https://www.youtube.com/watch?v=0704iLc55Fs
Look at the papers
- Function Vectors in large language models
- Emergent World Representations
- Designing a Dashboard for transparency and control of conversational AI
Website to that does an interpretability on sparse autoencoders
https://www.neuronpedia.org/gemma-scope#main
Slides https://drive.google.com/file/d/1DqbVcNj-AMAecVJ_dpInV3Sx4-vL5cgG/view
© 2026 bsybin