Ev

Evals Skills

Hacker News

Evals Skills

Hello HN I'm Rogerio, co-founder of LangWatch This past month we've completely changed the way we onboard new customers now on LangWatch, instead of giving them instructions on how to instrument, cookbooks or UIs to build evals, or docs on how to write Scenario agent simulation tests, we simply give them skills now, or ready to copy-and-paste prompts. This has reduced our onboarding time to only a few minutes, no more postponing evals because other priorities gets in the way. We have now skills for everything for managing your agent lifecycle: "Instrument my agent with open telemetry" "Write evaluations for my agent" "Write scenario tests and a CI pipeline for my agent" "Version my prompts" and even more targeted recipes "Check my agent doesn't give prescriptive advice" "Generate an evaluation dataset from my RAG knowledge base" "Test my CLI is well usable by other AI agents" Check out more on LangWatch Skills directory above and lmk what you think

Share card

Actual performance

4points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: agents, agent, new · Missing: mac, macos, cursor
95%95% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
63%63% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: lua, pipe, io · Missing: https docs, excited, just released
52%52% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: month, way · Missing: mobile apps, ios, personal
49%49% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
25%25% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
22%22% predicted probability of success on AppSumo, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
4%4% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Zi
Zine on LLM Evals44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Zine on LLM Evals

Hacker News1
Cl
Claude Code skills for building LLM evals50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Claude Code skills for building LLM evals

Hacker News2
We
We wrote a book on system evals68%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

We wrote a book on system evals

Hacker News3
Ht
Htmx Skills46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Htmx Skills

Hacker News2
SkillFade
SkillFade58%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

See which skills are fading before it's too late

Product Hunt+6
Mt
Mthds – Beyond skills: a typed DSL for executable AI methods50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Mthds – Beyond skills: a typed DSL for executable AI methods

Hacker News23
We
We wrote a book on LLM system evals with a bear and fox68%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

We wrote a book on LLM system evals with a bear and fox

Hacker News11
Sk
Skills.wtf – Find the Best Alexa Skills31%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Skills.wtf – Find the Best Alexa Skills

Hacker News1
BragSkills
BragSkills43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Brag about your skills online

Indie Hackers2b2b
Do
Donate your time/skills for a charitable donation45%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Donate your time/skills for a charitable donation

Hacker News5