Li

Litmus – Specification testing for structured LLM outputs

Hacker News

Litmus – Specification testing for structured LLM outputs

Over the holidays, I've been working on a small side-project that includes some LLM prompting from the end user. Admittedly, I struggle to keep track of the latest and greatest models, and I've also never bothered to read up on "prompt engineering," so I built a little testing utility to solve both of these problems at the same time. Enter Litmus. I'm pitching it as "specification testing" for LLMs. You define test cases (input prompt -> output JSON), as well as your system prompt and structured output (JSON Schema). All of this gets chucked at OpenRouter, and you get some nice terminal output summarising the test results (with a breakdown per-field for any failing cases) to see how well the model performed. Although it's framed as an LLM testing tool, it also serves as a model comparator. You can pass the `--model` CLI argument multiple times, and this will let you run the test cases against multiple models, with a comparison table generated in the output at the end for evaluating latency, throughput, tokens, and accuracy (tests passing vs. failing). The GitHub README contains a full example output of what a test report from Litmus looks like. With this, I've managed to get my system prompt for my side-project whittled down to the point where the accuracy is acceptable and it's not an exorbitant amount of tokens. I've also found out, through model comparison, that I didn't need anywhere near as large of a model as I had originally envisioned. You can grab it on GitHub as a single-file, zero-dependency executable (written in Go). Admittedly, I've not tested the pre-built binaries that are created via GitHub Actions, but there's no reason why they shouldn't work.

Share card

Actual performance

1points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model, user, models · Missing: mac, agents, macos
93%93% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: created, para · Missing: supports, reddit linkedin, podcasting
90%90% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: para · Missing: mobile apps, ios, personal
41%41% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: lua, ide, io · Missing: https docs, excited, just released
40%40% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
26%26% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
15%15% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Pg
Pg_yregress, Structured Testing for Postgres45%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Pg_yregress, Structured Testing for Postgres

Hacker News56
Af
AfricanaWiki – the structured Africana knowledgebase62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

AfricanaWiki – the structured Africana knowledgebase

Hacker News2
ty
typed_params – structured and typed parameters for Rails controllers44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

typed_params – structured and typed parameters for Rails controllers

Hacker News1
Xy
Xytext – Turn LLM Prompts into Structured API Endpoints61%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Xytext – Turn LLM Prompts into Structured API Endpoints

Hacker News2
LLM Knights
LLM Knights41%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The Unified Playground for LLM Testing

Indie Hackers1ai
La
Language-Agnostic Mutation Testing with LLM-Powered Agents34%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Language-Agnostic Mutation Testing with LLM-Powered Agents

Hacker News1
Fu
Fuzzpy, a fuzzer for testing CPython33%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Fuzzpy, a fuzzer for testing CPython

Hacker News6
Ch
ChefSpec – RSpec testing for Chef Cookbooks33%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

ChefSpec – RSpec testing for Chef Cookbooks

Hacker News3
A/
A/B testing baked into RequireJS33%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A/B testing baked into RequireJS

Hacker News1
A/
A/B testing in Golang with PlanOut30%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A/B testing in Golang with PlanOut

Hacker News8