Ca

Caliper – pass@k reliability testing for Claude Code and Codex skills

Hacker News

Caliper – pass@k reliability testing for Claude Code and Codex skills

Skills for Claude Code and Codex are hard to test. What I mean by hard is that there's no standard way to do it. You evaluate the skill once on something, it looks like it works. You publish it. Then the new super model releases (GLM 5.2 anyone?), it will quietly break for some part, and you won't find out until your users complain. I also faced the same problem, so I tried to build something lightweight to stop doing that. Caliper. It's a local and lightweight harness that runs a skill k times in isolated environments and gives you a pass@k score (How much times it succeeded in these k times). As a non-deterministic technology, you can't just say "it worked once". You need to answer how much it passed in k times. You define success in a YAML spec. I picked YAML to keep a schema and make it still readable for a human. You either use a LLM judge, a Python assertion, or both: Here's an simple evaluation example with a JSON extraction, so you write this in a YAML file: tasks: - name: Extracts action items as clean JSON prompt: "Read /tmp/transcript.txt and write the action items to /tmp/actions.json." expect: "A valid JSON array where every item has owner, task, due. No markdown fences." assert: | import json items = json.load(open("/tmp/actions.json")) assert isinstance(items, list) assert all({"owner","task","due"} <= i.keys() for i in items) Then with the CLI, you'll run it: caliper run extract-actions.eval.yaml --k 5 --baseline What's cool about the --baseline flag is that it will re-runs everything without the skill, so you can see whether the skill is doing the work or the base agent was going to pass anyway: ID Task k(5) pass@k task-1 Extracts action items as JSON 5/5 100% PASS With skill 100% No skill 60% Delta +40% Most models know how to get the JSON right most of the time (JSON extraction was solved by 2 years old already). But that's it, "most of the time" is the bug. That delta shows how the skill actually helped. (It's sometimes 0%, sometimes -100%!) I also created two skills you can get started right away with your favorite harness, e.g. Claude Code, Codex or Pi: - evaluate-skill: run and manage evals without leaving your workflow - grill-skill: reads your SKILL.md, interviews you about what "good" looks like, writes a 3-task spec (happy path, edge case, adversarial), and runs it You can install the skill with the command: npx skills@latest add edonadei/caliper I for now support claude-code, codex, pi, claude-api, openai-api. You can run the agent and the judge as separate backends, so you can run a skill on one and judge with another. GitHub: https://github.com/edonadei/caliper PyPI: https://pypi.org/project/caliper-eval/ Of course, it's a first step. I think the autorater layer can be vastly improved, more handholding to create and iterate on evaluation specs, supporting more harness, why not including this layer into a self-improvement bigger system? If you're also building agentic evaluations, I'm genuinely interested to hear how you are handling that.

Share card

Actual performance

3points
3comments
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: agent, claude, model · Missing: mac, agents, macos
96%96% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: created, started, para · Missing: supports, reddit linkedin, podcasting
87%87% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
TrustMRRFits verified-revenue profile · Strong signals: users, way, para · Missing: mobile apps, ios, personal
54%54% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: lua, io, including · Missing: https docs, excited, just released
26%26% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
Acquire.comPre-revenue stage for this audience · Strong signals: arr · Missing: mrr, revenue, profit
23%23% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: users · Missing: plus, platform, intuitive
21%21% predicted probability of success on AppSumo, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Co
Code Security Skills Codex-Inspired Workflows Packaged for Claude Code45%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Code Security Skills Codex-Inspired Workflows Packaged for Claude Code

Hacker News2
iT
iTerm2 Plugin for Codex/Claude Code30%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

iTerm2 Plugin for Codex/Claude Code

Hacker News2
Cl
Claude Code Skills Playground41%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Claude Code Skills Playground

Hacker News4
Ev
Ever Wanted to Call Codex from Claude Code? My Harness Orchestrator37%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Ever Wanted to Call Codex from Claude Code? My Harness Orchestrator

Hacker News3
Pr
Promode for Claude Code30%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Promode for Claude Code

Hacker News1
TA
TAS – Tracking, Automation, and Skills for Claude Code30%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

TAS – Tracking, Automation, and Skills for Claude Code

Hacker News1
Po
Poka-Yoke – mistake-proofing skills for Claude Code, with a benchmark34%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Poka-Yoke – mistake-proofing skills for Claude Code, with a benchmark

Hacker News1
HiveTechs Consensus
HiveTechs Consensus38%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

One workspace for Claude Code, Gemini, Codex, and 8 more

Indie Hackerscommitment-full-time
Bu
BugMagnet for Claude Code and Cursor – automated exploratory testing30%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

BugMagnet for Claude Code and Cursor – automated exploratory testing

Hacker News2
sk
skillhealth - find Claude Code skills you never use and what they cost39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

skillhealth - find Claude Code skills you never use and what they cost

Hacker News3