Llama 3.2 Interpretability with Sparse Autoencoders
Llama 3.2 Interpretability with Sparse Autoencoders
I spent a lot of time and money on this rather big side project of mine that attempts to replicate the mechanistic interpretability research on proprietary LLMs that was quite popular this year and produced great research papers by Anthropic [1], OpenAI [2] and Deepmind [3]. I am quite proud of this project and since I consider myself the target audience for HackerNews did I think that maybe some of you would appreciate this open research replication as well. Happy to answer any questions or face any feedback. Cheers [1] https://transformer-circuits.pub/2024/scaling-monosemanticit... [2] https://arxiv.org/abs/2406.04093 [3] https://arxiv.org/abs/2408.05147
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Correct prediction on native model
Similar products
Llama or Alpaca?
Llama 2 Uncensored 70B as API
Finetune Llama-3.1 2x faster in a Colab
Finetune Llama 3.2 Vision in a Colab
TokenHawk, WebGPU Running LLaMA
Llama Running on a Microcontroller
Llama list – todolists are better with a llama companion
Python Bindings for llama.cpp with some CLIs
Run Llama 3.1 8B in the browser
Llama 3.3 70B Sparse Autoencoders with API access