Sp

Speeding up LLM inference 2x times (possibly)

Hacker News

Speeding up LLM inference 2x times (possibly)

Here's a project I've been working on for the last few months. It's a new (I think) algorithm, that allows to adjust smoothly - and in real time - how many calculations you'd like to do during inference of an LLM model. It seems that it's possible to do just 20-25% of weight multiplications instead of all of them, and still get good inference results. I implemented it to run on M1/M2/M3 GPU. The mmul approximation itself can be pushed to run 2x fast before the quality of output collapses. The inference speed is just a bit faster than Llama.cpp's, because the rest of implementation could be better, but with a better development I think it can be a new method to speed up inference - in addition to quantization. You could call it ad-hoc model distillation :) You can change the speed / accuracy of a model at will, in real time. Oh, and as a side effect, the data format allows to also choose how much of the model you want to load into the memory. You can decide to skip say 10-20-40% of the least important weights. It's implemented for Mistral, it was also tested slightly on Mixtral and Llama. It's for FP16 for now, but Q8 is in the works. The algorithm is described here, and the implementation is open source. https://kolinko.github.io/effort/ I know these are bold claims, but I hope they survive the scrutiny :)

Share card

Actual performance

419points
114comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model, new, tiny · Missing: mac, agents, macos
91%91% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
89%89% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: open source, llama, ide · Missing: https docs, excited, just released
73%73% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRFits verified-revenue profile · Strong signals: month · Missing: mobile apps, ios, personal
52%52% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
30%30% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
16%16% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: real time · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

LL
LLM Inference Requirements Profiler59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LLM Inference Requirements Profiler

Hacker News4
Th
The Probability Times39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The Probability Times

Hacker News20
Ho
How many more times will you see your mother?40%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

How many more times will you see your mother?

Hacker News1
Op
Open-source AMDGCN kernels for optimizing LLM inference72%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-source AMDGCN kernels for optimizing LLM inference

Hacker News5
On
Onera – Private LLM Inference Inside AMD SEV-SNP Enclaves59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Onera – Private LLM Inference Inside AMD SEV-SNP Enclaves

Hacker News1
TimeCheck
TimeCheck22%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Find times when everyone is free

Indie Hackers1b2b
Bo
BonzAI – self-sovereign, local LLM inference in the browser62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

BonzAI – self-sovereign, local LLM inference in the browser

Hacker News5
En
Enfer.ai – Cheap LLM Inference Service58%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Enfer.ai – Cheap LLM Inference Service

Hacker News1
BroadBoard Times Square 🗽🌟
BroadBoard Times Square 🗽🌟46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Get your product launch featured in Times Square 🗽🌟

Indie Hackers1advertising
Ni
NightRun, bare metal LLM inference, no OS, boots from USB65%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

NightRun, bare metal LLM inference, no OS, boots from USB

Hacker News6