Mi

MinLlama – Llama 3.2 inference in ~100 lines of NumPy

Hacker News

MinLlama – Llama 3.2 inference in ~100 lines of NumPy

I built minLlama because I wanted a Llama implementation that was easy to understand and hack for KV cache compression research. There is also a PyTorch and Jax version in ~140 lines. Would be interested in feedback from people who have written transformer implementations before, are there any implementation "tricks" that I'm missing (e.g, cleaner KV cache for PyTorch/Jax or rope tricks)?

Share card

Actual performance

2points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Hacker NewsStrong engagement from HN community · Strong signals: llama, io · Missing: https docs, excited, just released
63%63% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
Product HuntOn track for Day 1 leaderboard · Missing: mac, agents, macos
59%59% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Indie HackersIH features products with proven revenue · Missing: supports, reddit linkedin, podcasting
40%40% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
35%35% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
32%32% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
24%24% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
2%2% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

A
A Brainfuck interpreter in 100 lines of C58%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A Brainfuck interpreter in 100 lines of C

Hacker News1
An
An ES6 Promises/A+ implementation in 100 lines of code44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

An ES6 Promises/A+ implementation in 100 lines of code

Hacker News1
Sn
SnappyBird in 100 lines of code50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

SnappyBird in 100 lines of code

Hacker News1
Re
Rebuilding GPT2 inference in ~500 lines of (commented) code68%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Rebuilding GPT2 inference in ~500 lines of (commented) code

Hacker News5
Ll
Llama 3.1 8B CPU Inference in a Browser via WebAssembly85%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Llama 3.1 8B CPU Inference in a Browser via WebAssembly

Hacker News4
Im
Implementing Unsure Calculator in 100 lines of Haskell57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Implementing Unsure Calculator in 100 lines of Haskell

Hacker News5
Ll
Llama or Alpaca?78%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Llama or Alpaca?

Hacker News6
Ll
Llama 3.2 Interpretability with Sparse Autoencoders74%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Llama 3.2 Interpretability with Sparse Autoencoders

Hacker News579
Jl
Jlama – A fast Java inference engine for GPT and Llama models69%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Jlama – A fast Java inference engine for GPT and Llama models

Hacker News7
Fl
Flappy Bird in 128 Lines of CoffeeScript43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Flappy Bird in 128 Lines of CoffeeScript

Hacker News6