A

A new engine to run Kimi K3 on a laptop

Hacker News

A new engine to run Kimi K3 on a laptop

Kimi K3 has 2.78 trillion parameters and ships as 1.42 TB of weights. It clearly does not fit in the memory of a laptop. But K3 is a Mixture-of-Experts model. For each token, only a small fraction of its 896 experts per layer is activated. That changes the problem: the entire model does not need to be resident in RAM, as long as the weights required by each token can be reached quickly enough. We built WASTE — the Weight-Aware Streaming Tensor Engine — to explore that idea. WASTE keeps the dense, repeatedly used part of the model resident in memory, stores the routed experts in an NVMe-optimized container, and streams only the experts selected during inference. The remaining RAM is used as a bounded expert cache. The current Kimi K3 container is 982 GiB. On a 64 GB MacBook Pro, WASTE runs the complete model at around 0.32–0.34 tokens per second, with a measured minimum memory requirement of approximately 29 GB at a 4K context. That is obviously not interactive performance yet. But the result we found interesting is that it works at all: this is the full open-weights model, not a distillation, a pruned version, or a smaller model using the Kimi name. The engine is written in C and has no BLAS, CUDA, ONNX, or Python dependency in the inference path. The same code can be used through the CLI, embedded as a library, or exposed through the included OpenAI-compatible server. Correctness was the first constraint. Every layer was validated against a PyTorch reference, with final logits matching within 3.6e-06. The vision tower is supported as well and matches its reference within 2.3e-06. The current bottleneck is understood: K3 needs roughly 17 GB of expert data per token, and more than half of the decode time is spent reading experts from disk. The engine is already operating close to the measured throughput limit of the laptop’s internal SSD. The next improvements therefore need to reduce the number of bytes read per token and increase useful expert reuse without pushing the operating system into paging. K3 is deliberately the extreme case. The same engine runs Kimi-Linear 48B from a 19 GB container at 8.92 tokens per second with an 8 GB memory budget. The broader goal is to make models that are much larger than available RAM usable locally, without sending private data to an API and without requiring specialized accelerator hardware. We have published the engine, container format, conversion tools, benchmarks, validation suite, and also the experiments that failed rather than quietly removing them. Everything is fully open source. Feedback on the storage layout, quantization, caching strategy, direct I/O, portability, and potential optimizations would be very welcome. Contributions of any kind — code, benchmarks, testing on different hardware, documentation, bug reports, or new ideas — are more than appreciated. Repo: https://github.com/sqliteai/waste

Share card

Actual performance

7points
3comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: mac, model, new · Missing: agents, macos, agent
95%95% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: para, compatible · Missing: supports, reddit linkedin, podcasting
92%92% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: open source, ide, io · Missing: https docs, excited, just released
65%65% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRFits verified-revenue profile · Strong signals: para · Missing: mobile apps, ios, personal
53%53% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
27%27% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: active · Missing: arr, mrr, revenue
19%19% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Ru
Run Full Kimi K3 with 29 GB of RAM70%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Run Full Kimi K3 with 29 GB of RAM

Hacker News9
To
Tom Nook's Laptop69%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Tom Nook's Laptop

Hacker News552
Pr
Proving 67M ZK rows on a laptop in 28s (Winterfell OOMs)63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Proving 67M ZK rows on a laptop in 28s (Winterfell OOMs)

Hacker News1
Laptop Bags
Laptop Bags26%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Handcrafted Leather laptop Bags

Indie Hackers
Wrapcart
Wrapcart36%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Wrapcart Lenovo LOQ Laptop

Indie Hackerscommitment-full-time
Laptop Oplader
Laptop Oplader50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Altijd de juiste laptop oplader

Product Hunt
Im
Implementation of Kimi K3 in PyTorch39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Implementation of Kimi K3 in PyTorch

Hacker News2
Ai
Aina, a templating engine for the new milenium45%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Aina, a templating engine for the new milenium

Hacker News1
Slap Your Laptop
Slap Your Laptop60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A stupidly fun laptop slapping toy.

Indie Hackers1$100/mogames
Wh
What Will You Build with Kimi K3?53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

What Will You Build with Kimi K3?

Hacker News2