CL

CLI for testing and evaluating LLM prompts and outputs

Hacker News

CLI for testing and evaluating LLM prompts and outputs

Hi HN, This project has grown a lot recently and figure it's worth another submission. I use this tool for several LLM-based use cases that have over 100k DAU. It works pretty simply: 1) Create a list of test cases 2) Set up assertions for metrics/guardrails you care about, such as outputting only JSON or not saying "As an AI language model" 3) Run tests as you make changes. Integrate with CI if desired. This makes LLM model and prompt selection easier because it reduces the process to something we're all familiar with: developing against test cases. You can iterate with confidence and avoid regressions. There are a bunch of startups popping up in this space, but I think it's important to have something that is local (private), on the command line (easy to use in the development loop), and open-source.

Share card

Actual performance

2points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model, open · Missing: mac, agents, macos
94%94% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
70%70% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: lua, ide, io · Missing: https docs, excited, just released
51%51% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
49%49% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
34%34% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
14%14% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
1%1% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Re
Recursive LLM Prompts53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Recursive LLM Prompts

Hacker News97
pr
promptmeter – LLM powered terminal load testing via prompts51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

promptmeter – LLM powered terminal load testing via prompts

Hacker News1
CL
CLI Testing Library35%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

CLI Testing Library

Hacker News2
Zs
Zsh Command Completion for Simon Wilson's LLM CLI54%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Zsh Command Completion for Simon Wilson's LLM CLI

Hacker News1
Li
Litmus – Specification testing for structured LLM outputs41%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Litmus – Specification testing for structured LLM outputs

Hacker News1
Er
Ergonomically call LLM in bulk from CLI50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Ergonomically call LLM in bulk from CLI

Hacker News7
TU
TUI Test – E2E testing for CLI / TUI apps44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

TUI Test – E2E testing for CLI / TUI apps

Hacker News3
Lo
LoadGQL – a CLI for load-testing GraphQL endpoints48%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LoadGQL – a CLI for load-testing GraphQL endpoints

Hacker News4
Pr
PrePrompt – rewrites vague prompts before they reach the LLM44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

PrePrompt – rewrites vague prompts before they reach the LLM

Hacker News23
LLM Knights
LLM Knights41%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The Unified Playground for LLM Testing

Indie Hackers1ai