Up

UpTrain (YC W23) – open-source tool to evaluate LLM response quality

Hacker News

UpTrain (YC W23) – open-source tool to evaluate LLM response quality

Hello, we are Shikha and Sourabh, founders of UpTrain(YC W23) - an open-source tool to evaluate the performance of your LLM applications on aspects such as correctness, tonality, hallucination, fluency, etc. The Problem: Unlike traditional Machine learning or Deep learning models where we always have a unique Ground Truth and can define metrics like Precision, Recall, accuracy, etc. to quantify the model’s performance, LLMs are trickier and it is very difficult to estimate if their response is correct or not. If you are using GPT-4 to write a recruitment email, there is no unique correct email to do a word-to-word comparison against. As you build an LLM application, you want to compare it against different model providers, prompt configurations, etc., and figure out the best working combination. Instead of manually skimming through a couple of model responses, you want to run them through hundreds of test cases, aggregate their scores, and make an informed decision. Additionally, as your application generates responses for real user queries, you don’t want to wait for them to complain about the model inaccuracy, instead, you want to monitor the model’s performance over time and get alerted in case of any drifts. Again, at the core of it, you want a tool to evaluate the quality of your LLM response and assign quantitative scores. The Solution: To solve this, we are building UpTrain which has a set of evaluation metrics so that you can know when your application is going wrong. These metrics include traditional NLP metrics like Rogue, Bleu, etc., embeddings similarity metrics as well as model grading scores i.e. where we use LLMs to evaluate different aspects of your response. A few of these evaluation metrics include: 1. Response Relevancy: Measures if the response contains any irrelevant information 2. Response Completeness: Measures if the response answers all aspects of the given question 3. Factual Accuracy: Measures hallucinations i.e. if the response has any made-up information or not with respect to the provided context 4. Retrieved Context Quality: Measures if the retrieved context has sufficient information to answer the given question 5. Response Tonality: Measures if the response aligns with a specific persona or desired tone etc. We have designed workflows so that you can easily add your testing dataset, configure which checks you want to run (you can also define custom checks suitable for your use case) and conveniently access the results via Streamlit dashboards. UpTrain also has experimentation capabilities where you can specify different prompt variations and models to test across and use these quantitative checks to find the best configuration for your application. You can also use UpTrain to monitor your application’s performance and find avenues for improvement. We integrate directly with your databases (BigQuery, Postgres, MongoDB, etc.) and can run daily evaluations. We’ve launched the tool under an Apache 2.0 license to make it easy for everyone to integrate it into their LLM workflows. Additionally, we also provide managed service (with a free trial) where you can run LLM evaluations via an API request or through UpTrain testing console. We would love for you to try it out and give your feedback. Links: Demo: https://demo.uptrain.ai/evals_demo/ Github repo: https://github.com/uptrain-ai/uptrain Create an account (free): https://uptrain.ai/dashboard UpTrain testing console (need an account): https://demo.uptrain.ai/dashboard Website: https://uptrain.ai/

Share card

Actual performance

12points
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: mac, model, user · Missing: agents, macos, agent
91%91% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
86%86% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: lua, ide, io · Missing: https docs, excited, just released
54%54% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: answers, way · Missing: mobile apps, ios, personal
48%48% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
39%39% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
10%10% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Me
Metlo (YC S21) – An Open Source API Security Tool81%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Metlo (YC S21) – An Open Source API Security Tool

Hacker News34
Ch
Chat with 4500 YC comapanies data and its open source61%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Chat with 4500 YC comapanies data and its open source

Hacker News1
We
Weaklayer – open-source Browser Detection and Response68%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Weaklayer – open-source Browser Detection and Response

Hacker News6
Fl
Flasho (YC W20) -Open source tool to set up transactional notifications69%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Flasho (YC W20) -Open source tool to set up transactional notifications

Hacker News2
In
InfraGenie – open-source tool to decouple your Terraform54%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

InfraGenie – open-source tool to decouple your Terraform

Hacker News3
Op
Open-Source Osint Tool71%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-Source Osint Tool

Hacker News1
K4
K4nundrum – An open-source tool to investigate Kryptos K462%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

K4nundrum – An open-source tool to investigate Kryptos K4

Hacker News1
Co
CodeOrb – Open-source µC debugging tool49%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

CodeOrb – Open-source µC debugging tool

Hacker News2
Sp
SpiderFoot, the open source footprinting tool71%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

SpiderFoot, the open source footprinting tool

Hacker News2
Sh
Shireframe – Open source declarative wireframing tool for programmers78%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Shireframe – Open source declarative wireframing tool for programmers

Hacker News34