In

InferShrink – Cut LLM API costs 10x with automatic model routing

Hacker News

InferShrink – Cut LLM API costs 10x with automatic model routing

I built this to solve my own problem — paying for GPT-4/Claude on prompts that Gemini Flash handles fine. InferShrink wraps your existing OpenAI/Anthropic/Google client in 3 lines. It classifies prompt complexity and routes to the cheapest model that can handle it. Same provider, no surprise switches. The pipeline: classify → compress (LLMLingua, optional) → retrieve (FAISS, optional) → route → track. When all stages combine, 10x+ cost reduction on mixed workloads. Key design decisions: • Same-provider routing only. If you use OpenAI, it stays on OpenAI. No cross-provider surprises. • Sub-millisecond classification overhead • Optional FAISS retrieval + LLMLingua compression for RAG pipelines • 539 tests, Semgrep + Trivy scanned pip install infershrink Blog post with the reasoning: https://musashimiyamoto1-cloud.github.io/infershrink-site/bl... Happy to answer questions about the routing heuristics or compression tradeoffs.

Share card

Actual performance

2points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: claude, model, google · Missing: mac, agents, macos
91%91% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: gemini · Missing: supports, reddit linkedin, podcasting
68%68% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: google · Missing: mobile apps, ios, personal
47%47% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: exist, existing, ide · Missing: https docs, excited, just released
45%45% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
38%38% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
17%17% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
3%3% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

LL
LLmHub.dev – A Unified API for Multi-LLM Automatic Routing51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LLmHub.dev – A Unified API for Multi-LLM Automatic Routing

Hacker News1
Wa
WatchLLM – Semantic caching to cut LLM API costs by 70%44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

WatchLLM – Semantic caching to cut LLM API costs by 70%

Hacker News1
I
I made a UI library with automatic routing and no “props” concept59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I made a UI library with automatic routing and no “props” concept

Hacker News2
Mo
Model-literals, model-aliases, and preference-aligned routing for LLMs42%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Model-literals, model-aliases, and preference-aligned routing for LLMs

Hacker News2
Sudo AI
Sudo AI83%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

One API for any LLM— routing, context, and monetization

Product Hunt+262API
KeksCode
KeksCode12%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Angular bits to cut your product development costs

Indie Hackers
Ad
Adaptive RAG – How we cut LLM costs without sacrificing accuracy38%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Adaptive RAG – How we cut LLM costs without sacrificing accuracy

Hacker News8
I
I built a crypto data API that ignores CEXs to cut costs by 85%33%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I built a crypto data API that ignores CEXs to cut costs by 85%

Hacker News2
An
An open-source UI Library with automatic routing and SuperComponents73%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

An open-source UI Library with automatic routing and SuperComponents

Hacker News7
Ro
RouteGPT – model routing on ChatGPT aligned to user preferences32%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

RouteGPT – model routing on ChatGPT aligned to user preferences

Hacker News2