pr

prompttest – pytest for LLMs

Hacker News

prompttest – pytest for LLMs

I’ve been experimenting with LLMs lately and kept running into the same problem: Every time I tweaked a prompt, I had to rerun a bunch of test cases manually and eyeball the results. It felt like writing code without unit tests. Existing tools I found were either: - Full frameworks where you write evaluators in Python. - Big platforms expanding into monitoring/security. I wanted something simpler: a fast CLI that just tests prompts. So I built prompttest, a pytest-like workflow for LLMs: - You define a prompt in a .txt file with `{variables}`. - You write test cases in .yml, with plain-English criteria. - You run prompttest to see clear pass/fail results in your terminal. The core idea is that your "assertion" is just English. Example: > The response must be polite and address the user by name. Then a model grades the output for you. This gives a safety net — you can refactor prompts and instantly see regressions. The project is still early. It runs on OpenRouter, so you can test against many models (including free ones) with one API key. Would love feedback, ideas, or use-cases you’d want supported. GitHub: https://github.com/decodingchris/prompttest

Share card

Actual performance

2points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model, user, models · Missing: mac, agents, macos
87%87% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: including · Missing: supports, reddit linkedin, podcasting
83%83% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: platform · Missing: plus, intuitive, reviews
43%43% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: exist, lua, existing · Missing: https docs, excited, just released
34%34% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
32%32% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
15%15% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

GP
GPTCache – Redis for LLMs69%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

GPTCache – Redis for LLMs

Hacker News7
Kr
KraspAI Kompass – keep up with new LLMs53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

KraspAI Kompass – keep up with new LLMs

Hacker News1
Integri
Integri42%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LLMs in one platform for free

Indie Hackerscommitment-side-project
Dageno AI
Dageno AI77%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Become the most recommended brand across 7+ major LLMs

Product Hunt+234Marketing
De
DeepTeam – Penetration Testing for LLMs51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

DeepTeam – Penetration Testing for LLMs

Hacker News3
PromptPassport
PromptPassport50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Compare All the LLMs at once

Product Hunt+8
Ru
Running LLMs on CPUs55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Running LLMs on CPUs

Hacker News1
GA
GAI, a Go-idiomatic, lightweight abstraction on top of LLMs52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

GAI, a Go-idiomatic, lightweight abstraction on top of LLMs

Hacker News2
Ja
Jax and Flax LLMs – Transformer Implementations Optimized for TPUs70%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Jax and Flax LLMs – Transformer Implementations Optimized for TPUs

Hacker News3
Ca
Call Multiple LLMs with GraphQL and AI Chainer42%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Call Multiple LLMs with GraphQL and AI Chainer

Hacker News2