WatchLLM – Semantic caching to cut LLM API costs by 70%
WatchLLM – Semantic caching to cut LLM API costs by 70%
Hey HN! I just shipped WatchLLM - a semantic caching layer for LLM APIs that sits between your app and providers like OpenAI/Claude/Groq. The problem: LLM API costs add up fast, especially when users ask similar questions in different ways ("how do I reset my password" vs "I forgot my password"). The solution: Semantic caching. WatchLLM vectorizes prompts, checks for similar queries (95%+ similarity), and returns cached responses instantly (50ms). If it's a miss, we forward to the actual API and cache for next time. Built in 3 days with Node.js, TypeScript, React, Cloudflare Workers (edge deployment), D1, and Redis. Just added prompt normalization today to boost cache hit rates even further. It's drop-in - literally just change your baseURL and keep using your existing OpenAI/Claude SDKs. No code changes needed. Currently in beta with a generous free tier (50K requests/month). Would love feedback from anyone building LLM apps - especially on the semantic similarity threshold and normalization strategies. Live demo on the site shows real-time cache hits and savings.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Correct prediction on native model
Similar products
SemanticCache – Save 70%+ on LLM API costs with semantic caching (Ruby)
InferShrink – Cut LLM API costs 10x with automatic model routing
AI-GATEWAY that cuts LLM API TOKEN costs by 40-70%.
Angular bits to cut your product development costs
SchemaVer for semantic versioning of schemas
Adaptive RAG – How we cut LLM costs without sacrificing accuracy
I built a crypto data API that ignores CEXs to cut costs by 85%
GoKubeDownscaler – Off-Hours Kubernetes Scaling Cuts Costs by 70%
Cut your AI token costs by 40-60% with one API call
Analyzing Semantic Redundancy in LLM Retrieval (Google GIST Protocol)