Cu

Cuts Long Horizon Inference Costs by 50% via external KV Cache Offload

Hacker News

Cuts Long Horizon Inference Costs by 50% via external KV Cache Offload

Hey HN, we’re the developers of OpenLake, an open source storage engine for offloading LLM KV caches from GPU memory into a shared tier of RAM and NVMe. We built OpenLake because KV caches are outgrowing GPU memory. A single 256K token conversation on Gemma 4 31B produces approximately 43GB of KV state, more than half the memory of an 80GB H100. The problem becomes even harder across a cluster: a prefix cached on one GPU host is unavailable when the next request lands on a different GPU, forcing the new GPU to repeat work the fleet has already completed. Once the KV cache is offloaded, network bandwidth becomes a major constraint on read latency. To move less data across the wire, we built deferred materialization: a custom CUDA kernel that losslessly compresses KV blocks before they leave GPU memory and decompresses them on the GPU after retrieval. In our tests, this achieved: - 1.72× lossless KV compression. - Approximately 600GB/s decompression throughput on an H100 - 80GB/s of effective KV throughput over a physical 50GB/s link At 128K context, retrieving cached KV reduces TTFT from 44 seconds to 0.6 seconds, a 66× improvement. Across the complete workload, GPU time reduces from 1,169 seconds to 606 seconds, saving 48.2% of GPU cost. OpenLake is written in Rust and uses io_uring with one pinned runtime per physical core. We provide connectors for vLLM and SGLang so the cache can be enabled without modifying the inference engine itself. I would love to hear how others are handling KV reuse across GPU hosts, especially for long contexts, and get to know your thoughts. Thanks! GitHub: https://github.com/openlake-project/openlake Here is our blog: https://cloud.theopenlake.com/blog/taming-the-beast-managing...

Share card

Actual performance

22points
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
92%92% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: new, context, single · Missing: mac, agents, macos
81%81% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: open source, ide, io · Missing: https docs, excited, just released
60%60% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Strong signals: host · Missing: plus, platform, intuitive
43%43% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
31%31% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
19%19% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Fi
Fig – Experimenting with long horizon prediction for personhood41%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Fig – Experimenting with long horizon prediction for personhood

Hacker News6
Op
OpenMetaHarness - complete long horizon tasks with more autonomy60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

OpenMetaHarness - complete long horizon tasks with more autonomy

Hacker News4
Te
Terminal-Bench-RL: Training long-horizon terminal agents with RL59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Terminal-Bench-RL: Training long-horizon terminal agents with RL

Hacker News125
Be
Benchmarking Tangible Interface Understanding in Long-Horizon Tasks51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Benchmarking Tangible Interface Understanding in Long-Horizon Tasks

Hacker News1
Se
Self-managing codebase with long-horizon agents57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Self-managing codebase with long-horizon agents

Hacker News2
Si
Single-agent long-horizon reasoning within one LLM run62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Single-agent long-horizon reasoning within one LLM run

Hacker News4
Op
OpenMetaHarness – long-horizon execution over multiple context sessions49%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

OpenMetaHarness – long-horizon execution over multiple context sessions

Hacker News3
Hy4 preview
Hy4 preview83%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Tencent’s 770B open model for long-horizon work

Product Hunt+216Open Source
Ci
CivBench a long-horizon AI benchmark for multi-agent games40%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

CivBench a long-horizon AI benchmark for multi-agent games

Hacker News12
Muse Code
Muse Code91%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Meta’s terminal agent for long-horizon coding

Product Hunt+244Productivity