KT

KTransformers–236B Model and 1M Context LLM Inference on Local Machines

Hacker News

KTransformers–236B Model and 1M Context LLM Inference on Local Machines

Hey Hacker News! We are excited to share our open-source project, KTransformers, a flexible framework designed for cutting-edge LLM inference optimizations! Leveraging state-of-the-art kernels from llamafile and marlin, KTransformers seamlessly enhances the performance of HuggingFace Transformers, making it possible to operate large 236B MoE models or extremely long 1M context locally with promising speed. KTransformers is a Python-centric framework designed with extensibility at its core. By implementing and injecting an optimized module with a single line of code, users gain access to a Transformers-compatible interface, RESTful APIs compliant with OpenAI and Ollama, and even a simplified ChatGPT-like web UI. For example, it allows you to integrate with all your familiar frontends, such as the VS Code plugin backed by Tabby. To demonstrate its capability, we present two showcase demos: - GPT-4-level Local VSCode Copilot: It runs the huge 236B DeepSeek-Coder-V2's Q4_K_M variant using just 11GB VRAM and 136GB DRAM on a local machine, which matches the score of GPT4-0613 in BigCodeBench with a promising 126 tokens/s for prompt prefill and 13.6 tokens/s for generation. - 1M Context Local Inference:Achieves 15 tokens/s with nearly 100% accuracy on the "Needle In a Haystack" test via the InternLM2.5-7B-Chat-1M model, utilizing 24GB VRAM and 150GB DRAM, and is several times faster than llama.cpp. Check it out on GitHub: https://github.com/kvcache-ai/ktransformers

Share card

Actual performance

20points
3comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: mac, model, user · Missing: agents, macos, agent
96%96% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: compatible · Missing: supports, reddit linkedin, podcasting
89%89% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: excited, llama, hacker news · Missing: https docs, just released, exist
66%66% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoStrong fit for a featured deal · Strong signals: interface, users · Missing: plus, platform, intuitive
61%61% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: users · Missing: mobile apps, ios, personal
47%47% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
13%13% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: chat · Missing: web3, crypto, cryptocurrency
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

LL
LLM Inference Requirements Profiler59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LLM Inference Requirements Profiler

Hacker News4
Bo
BonzAI – self-sovereign, local LLM inference in the browser62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

BonzAI – self-sovereign, local LLM inference in the browser

Hacker News5
We
WebGL 1M particles73%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

WebGL 1M particles

Hacker News8
Dr
Dragon – $1M Kitty65%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Dragon – $1M Kitty

Hacker News1
~1
~1m/Pixel EPSG 3857 Austrian LandCover Model57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

~1m/Pixel EPSG 3857 Austrian LandCover Model

Hacker News3
La
Launch StableStudio local inference in one commmand51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Launch StableStudio local inference in one commmand

Hacker News6
Ru
Run transformers model inference in C/C++ and even assembly60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Run transformers model inference in C/C++ and even assembly

Hacker News2
Ot
Otlet – Local LLM inference "inside" Postgres74%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Otlet – Local LLM inference "inside" Postgres

Hacker News2
Co
Collider – the platform for local LLM debug and inference at warp speed74%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Collider – the platform for local LLM debug and inference at warp speed

Hacker News3
Mu
Multi-agent LLM editor with local inference via WebSockets61%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Multi-agent LLM editor with local inference via WebSockets

Hacker News2