CLI for testing and evaluating LLM prompts and outputs
CLI for testing and evaluating LLM prompts and outputs
Hi HN, This project has grown a lot recently and figure it's worth another submission. I use this tool for several LLM-based use cases that have over 100k DAU. It works pretty simply: 1) Create a list of test cases 2) Set up assertions for metrics/guardrails you care about, such as outputting only JSON or not saying "As an AI language model" 3) Run tests as you make changes. Integrate with CI if desired. This makes LLM model and prompt selection easier because it reduces the process to something we're all familiar with: developing against test cases. You can iterate with confidence and avoid regressions. There are a bunch of startups popping up in this space, but I think it's important to have something that is local (private), on the command line (easy to use in the development loop), and open-source.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Incorrect prediction on native model
Similar products
Recursive LLM Prompts
promptmeter – LLM powered terminal load testing via prompts
CLI Testing Library
Zsh Command Completion for Simon Wilson's LLM CLI
Litmus – Specification testing for structured LLM outputs
Ergonomically call LLM in bulk from CLI
TUI Test – E2E testing for CLI / TUI apps
LoadGQL – a CLI for load-testing GraphQL endpoints
PrePrompt – rewrites vague prompts before they reach the LLM
The Unified Playground for LLM Testing