Op

Opik, an open source LLM evaluation framework

Hacker News

Opik, an open source LLM evaluation framework

Hey HN! I'm Caleb, one of the contributors to Opik, a new open source framework for LLM evaluations. Over the last few months, my colleagues and I have been working on a project to solve what we see as the most painful parts of writing evals for an LLM application. For this initial release, we've focused on a few core features that we think are the most essential: - Simplifying the implementation of more complex LLM-based evaluation metrics, like Hallucination and Moderation. - Enabling step-by-step tracking, such that you can test and debug each individual component of your LLM application, even in more complex multi-agent architectures. - Exposing an API for "model unit tests" (built on Pytest), to allow you to run evals as part of your CI/CD pipelines - Providing an easy UI for scoring, annotating, and versioning your logged LLM data, for further evaluation or training. It's often hard to feel like you can trust an LLM application in production, not just because of the stochastic nature of the model, but because of the opaqueness of the application itself. Our belief is that with better tooling for evaluations, we can meaningfully improve this situation, and unlock a new wave of LLM applications. You can run Opik locally, or with a free API key via our cloud platform. You can use it with any model server or hosted model, but we currently have a built-in integration with the OpenAI Python library, which means it automatically works not just with OpenAI models, but with any model served via a compatible model server (ollama, vLLM, etc). Opik also currently has out-of-the-box integrations with LangChain, LlamaIndex, Ragas, and a few other popular tools. This is our initial release of Opik, so if you have any feedback or questions, I'd love to hear them!

Share card

Actual performance

86points
15comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: agent, model, new · Missing: mac, agents, macos
97%97% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: compatible · Missing: supports, reddit linkedin, podcasting
87%87% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: lua, open source, llama · Missing: https docs, excited, just released
79%79% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: month · Missing: mobile apps, ios, personal
47%47% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: platform, host · Missing: plus, intuitive, reviews
40%40% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: training · Missing: arr, mrr, revenue
14%14% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Do
Dokimos – LLM evaluation framework for Java44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Dokimos – LLM evaluation framework for Java

Hacker News1
Ti
Tiledesk – Open-Source LLM Chatbot Framework43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Tiledesk – Open-Source LLM Chatbot Framework

Hacker News4
AI
AI PM Evaluation Framework (Open Source)58%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

AI PM Evaluation Framework (Open Source)

Hacker News2
Ev
Evvo – an open source framework for distributed evolutionary algorithms75%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Evvo – an open source framework for distributed evolutionary algorithms

Hacker News3
Me
MetricFlow – open-source metric framework81%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

MetricFlow – open-source metric framework

Hacker News98
Di
Dissect – An open source DFIR framework73%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Dissect – An open source DFIR framework

Hacker News8
RedBlue
RedBlue48%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-Source Hypervideo Framework

Indie Hackers1apis
ZenML
ZenML50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

extensible, open-source MLOps framework

Indie Hackers4$1/moai
Open-source evaluation framework for AI agents
Open-source evaluation framework for AI agents36%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

aligned to the OWASP Agentic Security Initiative (ASI)

Indie Hackers1ai
Be
Benchmax, a new open-source RL environment framework for LLM finetuning64%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Benchmax, a new open-source RL environment framework for LLM finetuning

Hacker News1