Be

BerriAI – Monitor Hallucinations in LLMs (Sentry for LLM Apps)

Hacker News

BerriAI – Monitor Hallucinations in LLMs (Sentry for LLM Apps)

Hi HN - Ishaan and Krrish here from BerriAI. We’ve built a hallucination monitoring tool for LLM Apps in production, that can instantly identify language mistranslations (responding to a user in the incorrect language) and inventing new information errors (answering from information not in the prompt). Live demo here : https://logs.berri.ai/ We served over 1m+ chatGPT queries with our initial ‘chat with your data’ app. However, we had no ability to tell how any of the technical changes we made (e.g. moving from llama index to our own retrieval/qa system) impacted our users in production. Berri is super easy to integrate into your system - we added it to our previous product with just 2 lines of code! It’s super early days and we’re looking for others like us - people in production - pushing changes but unsure if/how they’re actually solving issues / improving their system over time. Thanks for taking the time to read this, we’re really happy to be posting here :) Krrish and Ishaan

Share card

Actual performance

4points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: apps, user, new · Missing: mac, agents, macos
96%96% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
83%83% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: llama, ide, io · Missing: https docs, excited, just released
68%68% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRFits verified-revenue profile · Strong signals: apps, users · Missing: mobile apps, ios, personal
57%57% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: users · Missing: plus, platform, intuitive
31%31% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
14%14% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: chat · Missing: web3, crypto, cryptocurrency
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

In
Inspect Element for LLM Apps52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Inspect Element for LLM Apps

Hacker News1
Mo
Monitor LLMs Using LangChain and Infino47%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Monitor LLMs Using LangChain and Infino

Hacker News1
LM
LMM for LLMs – A mental model for building LLM apps62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LMM for LLMs – A mental model for building LLM apps

Hacker News6
GP
GPTCache – Redis for LLMs69%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

GPTCache – Redis for LLMs

Hacker News7
pr
prompttest – pytest for LLMs34%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

prompttest – pytest for LLMs

Hacker News2
LL
LLM Council – Run multiple LLMs with critique and consensus eval52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LLM Council – Run multiple LLMs with critique and consensus eval

Hacker News4
Sy
System monitor for OSX50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

System monitor for OSX

Hacker News1
SL
SLA Monitor by StatusGator50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

SLA Monitor by StatusGator

Hacker News1
Ma
Magellan VHDL Monitor40%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Magellan VHDL Monitor

Hacker News2
Pe
Per-monitor workspaces for Compiz 0.8.x51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Per-monitor workspaces for Compiz 0.8.x

Hacker News8