Ro

Route your prompts to the best LLM

Hacker News

Route your prompts to the best LLM

Hey HN, we've just finished building a dynamic router for LLMs, which takes each prompt and sends it to the most appropriate model and provider. We'd love to know what you think! Here is a quick(ish) screen-recroding explaining how it works: https://youtu.be/ZpY6SIkBosE Best results when training a custom router on your own prompt data: https://youtu.be/9JYqNbIEac0 The router balances user preferences for quality, speed and cost. The end result is higher quality and faster LLM responses at lower cost. The quality for each candidate LLM is predicted ahead of time using a neural scoring function, which is a BERT-like architecture conditioned on the prompt and a latent representation of the LLM being scored. The different LLMs are queried across the batch dimension, with the neural scoring architecture taking a single latent representation of the LLM as input per forward pass. This makes the scoring function very modular to query for different LLM combinations. It is trained in a supervised manner on several open LLM datasets, using GPT4 as a judge. The cost and speed data is taken from our live benchmarks, updated every few hours across all continents. The final "loss function" is a linear combination of quality, cost, inter-token-latency and time-to-first-token, with the user effectively scaling the weighting factors of this linear combination. Smaller LLMs are often good enough for simple prompts, but knowing exactly how and when they might break is difficult. Simple perturbations of the phrasing can cause smaller LLMs to fail catastrophically, making them hard to rely on. For example, Gemma-7B converts numbers to strings and returns the "largest" string when asking for the "largest" number in a set, but works fine when asking for the "highest" or "maximum". The router is able to learn these quirky distributions, and ensure that the smaller, cheaper and faster LLMs are only used when there is high confidence that they will get the answer correct. Pricing-wise, we charge the same rates as the backend providers we route to, without taking any margins. We also give $50 in free credits to all new signups. The router can be used off-the-shelf, or it can be trained directly on your own data for improved performance. What do people think? Could this be useful? Feedback of all kinds is welcome!

Share card

Actual performance

298points
126comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model, user, new · Missing: mac, agents, macos
82%82% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
80%80% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
46%46% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
40%40% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: ide, io · Missing: https docs, excited, just released
34%34% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
Acquire.comPre-revenue stage for this audience · Strong signals: margin, training, margins · Missing: arr, mrr, revenue
15%15% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Re
Recursive LLM Prompts53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Recursive LLM Prompts

Hacker News97
St
Standalone KeystoneJS route callbacks47%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Standalone KeystoneJS route callbacks

Hacker News1
Ro
Route LLM prompts to cheapest capable model – pydantic-AI and litellm49%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Route LLM prompts to cheapest capable model – pydantic-AI and litellm

Hacker News1
bb
bbus.in BMTC Bus Route Search59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

bbus.in BMTC Bus Route Search

Hacker News2
Re
RequestHub – route webhooks from one service to others50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

RequestHub – route webhooks from one service to others

Hacker News15
CampTarget
CampTarget49%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A route planner

Indie Hackers2community
Optiway
Optiway52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Route Planner

Indie Hackers1$40,000/mob2b
Pr
PrePrompt – rewrites vague prompts before they reach the LLM44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

PrePrompt – rewrites vague prompts before they reach the LLM

Hacker News23
Coralflavor
Coralflavor33%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Unfiltered LLM That Doesn't Reject Prompts

Indie Hackerscommitment-side-project
Se
Serve LLM prompts via CDN (free and no account allowed)56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Serve LLM prompts via CDN (free and no account allowed)

Hacker News1