Notes imported from my Obsidian daily track.
- An Introduction to Mechanistic Interpretability – Neel Nanda IASEAI 2025
- Analysing encoded concepts in transformer language models
- Current Open research
- Debating with More Persuasive LLMs Leads to More Truthful Answers
- Deep Learning Our Miraculous Year 19901991
- Deep Neural Network properties
- Differential Neural Computer
- EVERYTHING, EVERYWHERE, ALL AT ONCE IS MECHANISTIC INTERPRETABILITY IDENTIFIABLE
- Finding Manifolds With Bilinear Autoencoders
- Gao, Jun et al. “Representation Degeneration Problem in Training Natural Language Generation Models.
- Graph Rag
- How can AI help India
- Important Links
- Information Theory
- LLM - A survey
- LLM Self-Correction in Vision Language Action Models via Simulation-Driven Optimization
- Liner Representation hypothesis
- Maistros A greek LLM through knowledge distillation
- Manifold Hypothesis
- Mech Interp, Safety, Alignment Research Map
- Mnist Dataset
- NOT ALL LANGUAGE MODEL FEATURES ARE ONE-DIMENSIONALLY LINEAR
- Noticing The watcher - LLM infer surveillance from blocked feedback
- Paper Outline
- Proposal - Building Blocks of intelligence
- Regularization
- Sentence Transformers
- Softmax
- Sparse Autoencoder
- The Manifold Turn in Mechanistic Interpretability- A Review of Geo-metric Approaches to Understanding Neural Network Internals
- The Origins of Representation Manifolds in LargeLanguage Models
- Untitled
- Visualizing mnist dataset
- When Language Overwrites Vision Over Alignmentand Geometric Debiasing in Vision Language Models
- autoencoders
- cot-paper-reading
- manifold steering revels the shared geometry of neural network representation and behavior
- polysemantic
- vanishing gradients
- world model lecture
© 2026 bsybin