MinLlama – Llama 3.2 inference in ~100 lines of NumPy
MinLlama – Llama 3.2 inference in ~100 lines of NumPy
I built minLlama because I wanted a Llama implementation that was easy to understand and hack for KV cache compression research. There is also a PyTorch and Jax version in ~140 lines. Would be interested in feedback from people who have written transformer implementations before, are there any implementation "tricks" that I'm missing (e.g, cleaner KV cache for PyTorch/Jax or rope tricks)?
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Incorrect prediction on native model
Similar products
A Brainfuck interpreter in 100 lines of C
An ES6 Promises/A+ implementation in 100 lines of code
SnappyBird in 100 lines of code
Rebuilding GPT2 inference in ~500 lines of (commented) code
Llama 3.1 8B CPU Inference in a Browser via WebAssembly
Implementing Unsure Calculator in 100 lines of Haskell
Llama or Alpaca?
Llama 3.2 Interpretability with Sparse Autoencoders
Jlama – A fast Java inference engine for GPT and Llama models
Flappy Bird in 128 Lines of CoffeeScript