Op

OptiLLMBench – Test how inference optimization tricks scale up LLMs

Hacker News

OptiLLMBench – Test how inference optimization tricks scale up LLMs

OptiLLMBench is a new benchmark designed to evaluate how different inference optimization techniques (like ReRead, Chain-of-Thought, etc.) can improve LLM performance without any model changes or fine-tuning. To help understand real-world impact, I've included first results with Gemini 2.0 Flash: ReRead (RE2): +5% accuracy, +14% faster Chain-of-Thought Reflection: +5% boost Base performance: 51% The benchmark evaluates models on: Math word problems (GSM8K) Formal mathematics (MMLU Math) Logical reasoning (AQUA-RAT) Yes/no comprehension (BoolQ) The code works as a drop-in proxy - just point your OpenAI compatible endpoint to it and it'll apply the optimizations automatically. Dataset: https://huggingface.co/datasets/codelion/optillmbench Code: https://github.com/codelion/optillm Would love feedback from the HN community on additional optimization techniques to include or ways to improve the benchmark. Note: The dataset and proxy are completely open source and support any OpenAI API compatible endpoint.

Share card

Actual performance

2points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model, new, models · Missing: mac, agents, macos
94%94% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: gemini, compatible · Missing: supports, reddit linkedin, podcasting
80%80% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: lua, open source, io · Missing: https docs, excited, just released
56%56% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
44%44% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: way · Missing: mobile apps, ios, personal
44%44% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
11%11% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Ca
Calculate VRAM Requirements to Train/Inference with Your LLMs40%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Calculate VRAM Requirements to Train/Inference with Your LLMs

Hacker News1
Inference Engine by GMI Cloud
Inference Engine by GMI Cloud74%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Fast multimodal-native inference at scale

Product Hunt+180Developer Tools
Wh
WhitestormJS r11: modularity, optimization for webpack and more!45%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

WhitestormJS r11: modularity, optimization for webpack and more!

Hacker News1
Ge
Geopt – GEneric OPTimization by Genetically Evolved OPeration Trees42%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Geopt – GEneric OPTimization by Genetically Evolved OPeration Trees

Hacker News1
Ta
Tail Recursion Optimization for the JVM55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Tail Recursion Optimization for the JVM

Hacker News107
Li
Lizard Optimization59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Lizard Optimization

Hacker News1
TQNN
TQNN55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Fault-tolerant inference for noisy and imperfect data.

Indie Hackers1ai
Joule
Joule44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

AI Optimization

Indie Hackers1ai
We
We just launched MegaAI. It's a 4k30fps, 4W, 4TOPS inference powerhouse69%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

We just launched MegaAI. It's a 4k30fps, 4W, 4TOPS inference powerhouse

Hacker News3
GP
GPTCache – Redis for LLMs69%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

GPTCache – Redis for LLMs

Hacker News7