Mo

Monitor ChatGPT Hallucinations – With ChatGPT

Hacker News

Monitor ChatGPT Hallucinations – With ChatGPT

Got tired of manually parsing all my chatGPT logs. So I built a real-time hallucination detector for my logs in production. Now instead of manually trying to figure out which, of my hundreds of logs, were bad responses (invented new facts, refused to answer, etc.) I can just get chatGPT to flag them for me. How does it work? Bettershot aims to detect 3 things: Was the question relevant to the data (i.e. filter out questions like "how's the weather?" if the chatbot's purpose was to answer questions on the history of jeans) If relevant then, Did the model response invent new information when answering the question (i.e. information that was not in the prompt passed in) Did the model refuse to answer the question (e.g. "Sorry as an AI language model...") We do this by using chatgpt (currently gpt-3.5-turbo-16k) to evaluate each prompt-response pair 5 times, sampling the most frequent result (e.g. if it evaluated it to 'True' 4 times out of 5, then it's probably a good response). Check out the repo to know more https://github.com/ClerkieAI/bettershot

Share card

Actual performance

1points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model, new, chatgpt · Missing: mac, agents, macos
87%87% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
74%74% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
45%45% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: lua, io · Missing: https docs, excited, just released
35%35% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
34%34% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
14%14% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: chat · Missing: web3, crypto, cryptocurrency
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Ch
ChatGPT – The Memoir25%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

ChatGPT – The Memoir

Hacker News2
Ch
ChatGPT in Emacs43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

ChatGPT in Emacs

Hacker News23
St
StackOverflow for ChatGPT43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

StackOverflow for ChatGPT

Hacker News2
Th
This Is How ChatGPT Will Be Monetized27%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

This Is How ChatGPT Will Be Monetized

Hacker News9
Ch
ChatGPT Wrote a “Memoir”23%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

ChatGPT Wrote a “Memoir”

Hacker News2
Ch
ChatGPT Pitches OnlyBots Concept52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

ChatGPT Pitches OnlyBots Concept

Hacker News2
Ti
Tic Tac Toe Against ChatGPT27%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Tic Tac Toe Against ChatGPT

Hacker News1
Al
Alongside ChatGPT and DALL-E Emacs shells50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Alongside ChatGPT and DALL-E Emacs shells

Hacker News1
Fi
Filtir – Fixing ChatGPT Hallucinations42%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Filtir – Fixing ChatGPT Hallucinations

Hacker News2
Hy
Hypnotizing ChatGPT40%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Hypnotizing ChatGPT

Hacker News2