Do

Dokimos – LLM evaluation framework for Java

Hacker News

Dokimos – LLM evaluation framework for Java

I'm working on an open-source project dokimos, because every LLM eval framework I found was Python and TypeScript-only, but a lot of companies will be building LLM apps and AI agents with Java. Key features: - JUnit 5 integration for test-driven evals - Works with LangChain4j - Framework-agnostic - Supports custom evaluators and datasets GitHub: https://github.com/dokimos-dev/dokimos Would love contributions or to team up with anyone who has Java experience and wants to work on this together.

Share card

Actual performance

1points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: agents, agent, apps · Missing: mac, macos, cursor
87%87% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: supports · Missing: reddit linkedin, podcasting, created
69%69% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: lua, io · Missing: https docs, excited, just released
44%44% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: apps · Missing: mobile apps, ios, personal
40%40% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
29%29% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
19%19% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Op
Opik, an open source LLM evaluation framework79%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Opik, an open source LLM evaluation framework

Hacker News86
se
seqeval - a Python framework for sequence labeling evaluation52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

seqeval - a Python framework for sequence labeling evaluation

Hacker News2
ai
aiide – A pragmatic framework to build LLM Co-pilots55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

aiide – A pragmatic framework to build LLM Co-pilots

Hacker News2
Ru
Rues an Expression Evaluation Sidecar56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Rues an Expression Evaluation Sidecar

Hacker News1
Ze
Zeno, an Interactive ML Evaluation Framework64%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Zeno, an Interactive ML Evaluation Framework

Hacker News2
Evalentum
Evalentum62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Evaluation Framework for Startups Ecosystem

Indie Hackerscommitment-full-time
Fa
Faster LLM evaluation with Bayesian optimization50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Faster LLM evaluation with Bayesian optimization

Hacker News131
Ha
Hallu – a web framework where an LLM hallucinates your app55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Hallu – a web framework where an LLM hallucinates your app

Hacker News2
A
A Multichannel LLM Chatbot framework in Java35%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A Multichannel LLM Chatbot framework in Java

Hacker News2
Op
Open Evaluation53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open Evaluation

Hacker News3