Ru

RunNburn – Run a 295B Moe from a 98GB GGUF on a 64GB RAM Desktop

Hacker News

RunNburn – Run a 295B Moe from a 98GB GGUF on a 64GB RAM Desktop

runNburn is an Apache-2.0 Rust inference engine for quantized GGUF models that are too big for your fast memory. The core idea: weights stay file-backed (mmap), host residency stays under an explicit byte budget (--ram-budget), and GPU caches are sized from detected free/total VRAM — never from device-name presets. There is no conversion step, no sidecar cache files, no silent requantization. The GGUF on disk is the single source of truth. The result that made me want to post this: Tencent's Hy3 (295B total / 21B active sparse MoE, a single 97.8 GiB Q2_K GGUF) runs on my desktop with 64 GB of RAM and one consumer NVIDIA GPU. The file is larger than RAM and VRAM combined; the selected experts for each token are pulled on demand (the newest path batches O_DIRECT reads through io_uring), while the pretrained routing is left untouched. On the same machine, same prompt, same decode length, a warm-run median gave ~5.5 tok/s decode vs ~2.0 tok/s for llama.cpp. To be upfront about scope: for models that fit comfortably in VRAM, llama.cpp is still faster than runNburn today — its CUDA kernels have years of tuning and we measure against it honestly (interleaved A/B runs, medians, and any "speedup" that changes output quality is rejected). runNburn's lane is the model that doesn't fit. What's in the box: - CLI, interactive chat, and an OpenAI-compatible server (chat/completions + responses + conversations, SSE streaming, stateful continuation with KV/SSM snapshot reuse). It's built as a single-owner personal server — one active generation is the optimization unit; continuous batching and multi-tenant throughput are explicit non-goals. - Architecture-aware paths: Llama family, Phi, Gemma, Qwen dense/hybrid/MoE (including GatedDeltaNet layers), Nemotron-H MoE, Hy3, GLM — plus in-model multi-token prediction (self-speculative decoding) with device-side verification where the GGUF ships a drafter. - Backends: CPU is the default (x86 AVX2, ARM NEON), CUDA and Metal are active, Vulkan/OpenCL are experimental. Android works through a small C ABI (rnb.h). - Native quantized kernels for Q2_K–Q6_K, Q4_0, Q8_0 — including the low-bit K-quants that big-MoE builds actually ship in. It's pre-1.0 and rough in places; recognition of an architecture doesn't mean every community variant works. But if you've got a model file bigger than your machine and you'd rather it run slowly than not at all, that's exactly the case it was built for. Happy to answer questions about the offloading design, the expert-streaming path, or the measurement protocol.

Share card

Actual performance

11points
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: mac, model, new · Missing: agents, macos, agent
97%97% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: including, compatible · Missing: supports, reddit linkedin, podcasting
95%95% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: llama, ide, io · Missing: https docs, excited, just released
69%69% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRFits verified-revenue profile · Strong signals: personal · Missing: mobile apps, ios, entrepreneurs
62%62% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: plus, host · Missing: platform, intuitive, reviews
48%48% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: active · Missing: arr, mrr, revenue
23%23% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: chat · Missing: web3, crypto, cryptocurrency
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Mo
Moe-Direct – MoE Models far larger than your RAM, on a consumer desktop53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Moe-Direct – MoE Models far larger than your RAM, on a consumer desktop

Hacker News1
Ro
Robinhood on Desktop54%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Robinhood on Desktop

Hacker News7
Ro
Robinhood Desktop54%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Robinhood Desktop

Hacker News19
Co
Containerized Xorg Desktop Accessed via SPICE/HTML564%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Containerized Xorg Desktop Accessed via SPICE/HTML5

Hacker News4
We
WebTorrent Desktop 0.22.051%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

WebTorrent Desktop 0.22.0

Hacker News2
OS
OS108 Preconfigured NetBSD Desktop54%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

OS108 Preconfigured NetBSD Desktop

Hacker News1
my
my desktop extender is now on Softpedia54%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

my desktop extender is now on Softpedia

Hacker News1
Wa
Wallpapper Splitter for Many Desktop51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Wallpapper Splitter for Many Desktop

Hacker News2
I
I Reinvented the iPod but for Desktop47%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I Reinvented the iPod but for Desktop

Hacker News6
Bo
Bowery Desktop54%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Bowery Desktop

Hacker News3