Br

Braintrust – Eval platform for AI products

Hacker News

Braintrust – Eval platform for AI products

Hey HN, We're excited to introduce Braintrust, a platform for running and tracking AI evaluations (“evals”) [1]. At my previous startup Impira and leading AI at Figma, we had this recurring problem where we never knew if changes we made to our products would improve or regress key user scenarios. We built some tooling to solve this problem and after talking to other developers learned that it was a widespread issue. Specifically, it’s challenging to establish a great dev loop that lets you systematically improve and ship high quality AI products. We worked with the teams at Zapier, Coda, and Replit to refine Braintrust. We consistently heard that they were facing challenges with evaluation, so we built Braintrust to help them. Today we’re releasing the product for anyone to use — including a free plan [3]. There are a lot of LLM tooling products on the market. Here are a few ways Braintrust is different: - Rather than showing eval metrics in your observability tools, Braintrust offers an “experiment tracking” workflow, meaning you can try out changes while developing them, and drill down into diffs between other experiments and git branches before you ship. Check out our docs [2] for more details. - We believe strongly that you should own your data and support on-premises and private VPC deployments. - We natively and equally support Typescript and Python. - We have a flexible free plan for builders and an unlimited free plan for academic and non-commercial open source projects [3]. We will introduce a self-service paid (”Pro”) tier, hopefully with feedback from this community. Our mission is to enable developers to build high quality, reliable AI products. We couldn’t be doing this without Elad Gil, who helped me incubate the initial idea and team, which today includes founding designer, Coleen Baik, and founding engineer, Manu Goyal. Also big thanks to David Song, from Elad’s team, who is also helping us. We’re excited to launch today [4], but we know there’s a lot left to build and are excited to hear your feedback. [1] https://www.braintrustdata.com [2] http://www.braintrustdata.com/docs/guides/evals [3] https://www.braintrustdata.com/pricing [4] https://www.braintrustdata.com/blog/reliable-ai

Share card

Actual performance

8points
2comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Strong signals: ios, including · Missing: supports, reddit linkedin, podcasting
92%92% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: user, new, open · Missing: mac, agents, macos
87%87% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: excited, lua, open source · Missing: https docs, just released, exist
78%78% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Strong signals: platform, builder · Missing: plus, intuitive, reviews
44%44% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: ios, way · Missing: mobile apps, personal, entrepreneurs
31%31% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: recurring · Missing: arr, mrr, revenue
14%14% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: paid, introduce · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Ra
Raindrop – Sentry for AI Products30%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Raindrop – Sentry for AI Products

Hacker News11
Raindrop
Raindrop60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Sentry for AI Products

Product Hunt+241Developer Tools
Gl
GladAItor – Judge AI Products for Free24%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

GladAItor – Judge AI Products for Free

Hacker News4
Da
Dawn – Analytics for AI Products33%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Dawn – Analytics for AI Products

Hacker News8
Skope
Skope73%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The billing system for AI products.

Product Hunt+132SaaS
St
Startcrowd – Collaborate on ambitious AI products30%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Startcrowd – Collaborate on ambitious AI products

Hacker News2
Meteron AI
Meteron AI47%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Rapidly build AI products

Indie Hackerscommitment-side-project
Laminar
Laminar80%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-source all-in-one platform for engineering AI products

Product Hunt+351Productivity
AIMPLabs
AIMPLabs51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A clean, brandable name for AI products and labs

Indie Hackers
Au
Autoblocks Annotate – data annotation platform for AI products62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Autoblocks Annotate – data annotation platform for AI products

Hacker News2