ZS

ZSE – Open-source LLM inference engine with 3.9s cold starts

Hacker News

ZSE – Open-source LLM inference engine with 3.9s cold starts

I've been building ZSE (Z Server Engine) for the past few weeks — an open-source LLM inference engine focused on two things nobody has fully solved together: memory efficiency and fast cold starts. The problem I was trying to solve: Running a 32B model normally requires ~64 GB VRAM. Most developers don't have that. And even when quantization helps with memory, cold starts with bitsandbytes NF4 take 2+ minutes on first load and 45–120 seconds on warm restarts — which kills serverless and autoscaling use cases. What ZSE does differently: Fits 32B in 19.3 GB VRAM (70% reduction vs FP16) — runs on a single A100-40GB Fits 7B in 5.2 GB VRAM (63% reduction) — runs on consumer GPUs Native .zse pre-quantized format with memory-mapped weights: 3.9s cold start for 7B, 21.4s for 32B — vs 45s and 120s with bitsandbytes, ~30s for vLLM All benchmarks verified on Modal A100-80GB (Feb 2026) It ships with: OpenAI-compatible API server (drop-in replacement) Interactive CLI (zse serve, zse chat, zse convert, zse hardware) Web dashboard with real-time GPU monitoring Continuous batching (3.45× throughput) GGUF support via llama.cpp CPU fallback — works without a GPU Rate limiting, audit logging, API key auth Install: ----- pip install zllm-zse zse serve Qwen/Qwen2.5-7B-Instruct For fast cold starts (one-time conversion): ----- zse convert Qwen/Qwen2.5-Coder-7B-Instruct -o qwen-7b.zse zse serve qwen-7b.zse # 3.9s every time The cold start improvement comes from the .zse format storing pre-quantized weights as memory-mapped safetensors — no quantization step at load time, no weight conversion, just mmap + GPU transfer. On NVMe SSDs this gets under 4 seconds for 7B. On spinning HDDs it'll be slower. All code is real — no mock implementations. Built at Zyora Labs. Apache 2.0. Happy to answer questions about the quantization approach, the .zse format design, or the memory efficiency techniques.

Share card

Actual performance

58points
9comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Strong signals: compatible · Missing: supports, reddit linkedin, podcasting
93%93% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: model, openai, single · Missing: mac, agents, macos
90%90% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: llama, io · Missing: https docs, excited, just released
73%73% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
45%45% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
27%27% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: active · Missing: arr, mrr, revenue
25%25% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: chat · Missing: web3, crypto, cryptocurrency
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Op
Open-source AMDGCN kernels for optimizing LLM inference72%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-source AMDGCN kernels for optimizing LLM inference

Hacker News5
Ne
Neuropod – Uber ATG's open source deep learning inference engine83%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Neuropod – Uber ATG's open source deep learning inference engine

Hacker News80
Qa
Qake WebGL Voxel Engine released open source75%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Qake WebGL Voxel Engine released open source

Hacker News5
io
ioquake3 – An Open-Source Quake 3 Engine72%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

ioquake3 – An Open-Source Quake 3 Engine

Hacker News1
LL
LLM Inference Requirements Profiler59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LLM Inference Requirements Profiler

Hacker News4
Ji
Jinfer – AI inference engine for the JVM. AI in a jar57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Jinfer – AI inference engine for the JVM. AI in a jar

Hacker News3
Be
Beehive – An open source IFTTT powered by Go's templating engine66%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Beehive – An open source IFTTT powered by Go's templating engine

Hacker News247
Op
Open-source web voxel engine80%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-source web voxel engine

Hacker News4
Op
Open source machine learning inference accelerators on FPGA73%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open source machine learning inference accelerators on FPGA

Hacker News57
Op
Open-source M.U.G.E.N fighting engine using Godot75%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-source M.U.G.E.N fighting engine using Godot

Hacker News2