Wa

WatchLLM – Semantic caching to cut LLM API costs by 70%

Hacker News

WatchLLM – Semantic caching to cut LLM API costs by 70%

Hey HN! I just shipped WatchLLM - a semantic caching layer for LLM APIs that sits between your app and providers like OpenAI/Claude/Groq. The problem: LLM API costs add up fast, especially when users ask similar questions in different ways ("how do I reset my password" vs "I forgot my password"). The solution: Semantic caching. WatchLLM vectorizes prompts, checks for similar queries (95%+ similarity), and returns cached responses instantly (50ms). If it's a miss, we forward to the actual API and cache for next time. Built in 3 days with Node.js, TypeScript, React, Cloudflare Workers (edge deployment), D1, and Redis. Just added prompt normalization today to boost cache hit rates even further. It's drop-in - literally just change your baseURL and keep using your existing OpenAI/Claude SDKs. No code changes needed. Currently in beta with a generous free tier (50K requests/month). Would love feedback from anyone building LLM apps - especially on the semantic similarity threshold and normalization strategies. Live demo on the site shows real-time cache hits and savings.

Share card

Actual performance

1points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: claude, apps, user · Missing: mac, agents, macos
91%91% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
73%73% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
TrustMRRFits verified-revenue profile · Strong signals: apps, month, users · Missing: mobile apps, ios, personal
51%51% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: exist, existing, ide · Missing: https docs, excited, just released
43%43% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Strong signals: users · Missing: plus, platform, intuitive
39%39% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
19%19% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Se
SemanticCache – Save 70%+ on LLM API costs with semantic caching (Ruby)50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

SemanticCache – Save 70%+ on LLM API costs with semantic caching (Ruby)

Hacker News2
In
InferShrink – Cut LLM API costs 10x with automatic model routing43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

InferShrink – Cut LLM API costs 10x with automatic model routing

Hacker News2
Aurko
Aurko52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

AI-GATEWAY that cuts LLM API TOKEN costs by 40-70%.

Indie Hackerscommitment-full-time
KeksCode
KeksCode12%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Angular bits to cut your product development costs

Indie Hackers
Sc
SchemaVer for semantic versioning of schemas64%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

SchemaVer for semantic versioning of schemas

Hacker News1
Ad
Adaptive RAG – How we cut LLM costs without sacrificing accuracy38%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Adaptive RAG – How we cut LLM costs without sacrificing accuracy

Hacker News8
I
I built a crypto data API that ignores CEXs to cut costs by 85%33%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I built a crypto data API that ignores CEXs to cut costs by 85%

Hacker News2
Go
GoKubeDownscaler – Off-Hours Kubernetes Scaling Cuts Costs by 70%57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

GoKubeDownscaler – Off-Hours Kubernetes Scaling Cuts Costs by 70%

Hacker News3
AgentReady
AgentReady55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Cut your AI token costs by 40-60% with one API call

Product Hunt+117API
An
Analyzing Semantic Redundancy in LLM Retrieval (Google GIST Protocol)45%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Analyzing Semantic Redundancy in LLM Retrieval (Google GIST Protocol)

Hacker News6