InferShrink – Cut LLM API costs 10x with automatic model routing
InferShrink – Cut LLM API costs 10x with automatic model routing
I built this to solve my own problem — paying for GPT-4/Claude on prompts that Gemini Flash handles fine. InferShrink wraps your existing OpenAI/Anthropic/Google client in 3 lines. It classifies prompt complexity and routes to the cheapest model that can handle it. Same provider, no surprise switches. The pipeline: classify → compress (LLMLingua, optional) → retrieve (FAISS, optional) → route → track. When all stages combine, 10x+ cost reduction on mixed workloads. Key design decisions: • Same-provider routing only. If you use OpenAI, it stays on OpenAI. No cross-provider surprises. • Sub-millisecond classification overhead • Optional FAISS retrieval + LLMLingua compression for RAG pipelines • 539 tests, Semgrep + Trivy scanned pip install infershrink Blog post with the reasoning: https://musashimiyamoto1-cloud.github.io/infershrink-site/bl... Happy to answer questions about the routing heuristics or compression tradeoffs.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Correct prediction on native model
Similar products
LLmHub.dev – A Unified API for Multi-LLM Automatic Routing
WatchLLM – Semantic caching to cut LLM API costs by 70%
I made a UI library with automatic routing and no “props” concept
Model-literals, model-aliases, and preference-aligned routing for LLMs
One API for any LLM— routing, context, and monetization
Angular bits to cut your product development costs
Adaptive RAG – How we cut LLM costs without sacrificing accuracy
I built a crypto data API that ignores CEXs to cut costs by 85%
An open-source UI Library with automatic routing and SuperComponents
RouteGPT – model routing on ChatGPT aligned to user preferences