Ma

Mandoline – Custom LLM Evaluations for Real-World Use Cases

Hacker News

Mandoline – Custom LLM Evaluations for Real-World Use Cases

Hi HN! We're a small team of AI engineers who've spent the last few years building AI applications. Through this, we've experienced firsthand many of the challenges that come with evaluating and improving AI systems in real-world contexts. Standard LLM evaluations (and evaluation methods) often use simplified scenarios that don't reflect the complexity LLMs encounter in actual use. This leads to a disconnect between reported performance and real-world usefulness. We built Mandoline to bridge this gap, helping developers evaluate and improve LLM applications in ways that matter to end-users. Our approach allows you to design custom evaluation criteria that align with your specific product requirements. For a quick overview of how it works, check out our Python and / or Node SDK READMEs: - Python: https://github.com/mandoline-ai/mandoline-python/blob/main/R... - Node / Typescript: https://github.com/mandoline-ai/mandoline-node/blob/main/REA... Hopefully this design is flexible yet scalable, and helps you do things like: track LLM progress over time, make informed AI system design decisions, choose the best model for your use case, prompt engineer more systematically, and so on. Under the hood, Mandoline is a hybrid system using a combination of our own models and top general-purpose LLM APIs. We used Mandoline to evaluate and improve itself, which helped us make better decisions about system design. In the future, we’ll be adding visualization tools to more easily analyze trends, and expanding our in-house models capabilities to reduce reliance on (and hopefully outperform) external models. Check out our website ( https://mandoline.ai/ ) and documentation ( https://mandoline.ai/docs ) to learn more. We’d love to hear about your experiences with evaluating AI systems for production use. What have you found most challenging in evaluating AI systems? What behaviors are hard to quantify? How could Mandoline fit into your workflow? You can reach us here in the comments or send us an email (hi@mandoline.ai). We appreciate you taking the time to learn a bit about Mandoline.

Share card

Actual performance

2points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model, user, models · Missing: mac, agents, macos
90%90% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: ios · Missing: supports, reddit linkedin, podcasting
90%90% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
AppSumoStrong fit for a featured deal · Strong signals: users · Missing: plus, platform, intuitive
58%58% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRFits verified-revenue profile · Strong signals: ios, users, way · Missing: mobile apps, personal, entrepreneurs
56%56% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: lua, io · Missing: https docs, excited, just released
51%51% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
11%11% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Be
Benchmarking LLM Agents on Consequential Real World Tasks40%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Benchmarking LLM Agents on Consequential Real World Tasks

Hacker News3
Pa
PagerDuty for the real world48%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

PagerDuty for the real world

Hacker News1
I
I made Pokémon but with real animals in the real world62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I made Pokémon but with real animals in the real world

Hacker News4
Tr
Trying out actioncable in a real world app37%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Trying out actioncable in a real world app

Hacker News1
Cr
Crowsnest – API for the Real World52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Crowsnest – API for the Real World

Hacker News61
PharmaSafe
PharmaSafe52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Pharmacovigilance analytics for exploring real-world adverse

Indie Hackers1education
Th
The Whicher: A/B test the Real World40%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The Whicher: A/B test the Real World

Hacker News20
NL
NLP algorithms for real-world sentiment analysis48%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

NLP algorithms for real-world sentiment analysis

Hacker News1
Sethco AI
Sethco AI66%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Prepares your team for challenging real-world scenarios

Indie Hackers1$1,200/moai
MemoRep
MemoRep62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Spaced repetition for real-world skills.

Indie Hackers1education