LL

LLM Inference Performance Analytic Tool for Moe Models (DeepSeek/etc.)

Hacker News

LLM Inference Performance Analytic Tool for Moe Models (DeepSeek/etc.)

I built this to answer "what-if" questions about LLM deployment without spinning up expensive infrastructure. The tool models inference physics - latency, bandwidth saturation, and PCIe bottlenecks for large MoE models like DeepSeek-V3 (671B), Mixtral 8x7B, Qwen2.5-MoE, and Grok-1. Key features: - Independent Prefill vs Decode parallelism config (TP/PP/SP/DP) - Hardware modeling: H100, B200, A100, NVLink topologies, IB vs RoCE - Optimizations: Paged KV Cache, DualPipe, FP8/INT4 quantization - Experimental: Memory Pooling (TPP, tiered storage) and Near-Memory Computing - offload cold experts and cold/warm KV-cache to system RAM, node-shared or global-shared memory pool Live demo: https://llm-inference-performance-calculator-1066033662468.u... Built with React, TypeScript, Tailwind, and Vite. Disclaimer: I've calibrated the math models but they're not perfect. Feedback and PRs welcome.

Share card

Actual performance

1points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model, models, grok · Missing: mac, agents, macos
76%76% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: para · Missing: supports, reddit linkedin, podcasting
57%57% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
AppSumoStrong fit for a featured deal · Missing: plus, platform, intuitive
54%54% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: pipe, io · Missing: https docs, excited, just released
42%42% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: para · Missing: mobile apps, ios, personal
32%32% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
16%16% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
1%1% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

LL
LLM Inference Requirements Profiler59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LLM Inference Requirements Profiler

Hacker News4
YP
YPerf – Monitor LLM Inference API Performance58%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

YPerf – Monitor LLM Inference API Performance

Hacker News2
Ce
Cellulose – a tool to improve inference performance of ML models69%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Cellulose – a tool to improve inference performance of ML models

Hacker News11
Sp
Speeding up LLM inference 2x times (possibly)74%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Speeding up LLM inference 2x times (possibly)

Hacker News419
Op
Open-source AMDGCN kernels for optimizing LLM inference72%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-source AMDGCN kernels for optimizing LLM inference

Hacker News5
On
Onera – Private LLM Inference Inside AMD SEV-SNP Enclaves59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Onera – Private LLM Inference Inside AMD SEV-SNP Enclaves

Hacker News1
I
I made a tiny MoE/Engram viz tool74%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I made a tiny MoE/Engram viz tool

Hacker News2
Bo
BonzAI – self-sovereign, local LLM inference in the browser62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

BonzAI – self-sovereign, local LLM inference in the browser

Hacker News5
Op
Optimizing DeepSeek's NSA for TPUs – A Kernel Worklog48%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Optimizing DeepSeek's NSA for TPUs – A Kernel Worklog

Hacker News2
Mo
Moe-Direct – MoE Models far larger than your RAM, on a consumer desktop53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Moe-Direct – MoE Models far larger than your RAM, on a consumer desktop

Hacker News1