Sp

Spec27 – Spec-driven validation for AI agents

Hacker News

Spec27 – Spec-driven validation for AI agents

Hi HN! We’re a team of ML validation specialists and we’ve been building /Spec27, a tool for testing whether AI agents still do their job safely and reliably as models, prompts, tools, and surrounding systems change. We started working on this because a lot of current LLM evaluation work seems aimed at scoring general model behavior, while many teams are deploying systems that have a specific mission to fulfill. Many of the tools also assume you have full access to the agent stack and traces so you can place SDKs and Gateways, but a lot of agents are being created on vendor platforms where this isn’t possible. As a result, we approaches it from the outside in: all tests just run to the primary interfaces of an Agent and don’t assume anything about internals. The other important things about the approach is spec-driven. Instead of treating testing as a one-off benchmark or static eval set, we let teams define reusable specifications for the behavior they want from an agent, then generate tests against those specs. With this you can automatically generate adversarial and robustness checks, so you can see what an agent is sensitive to and what kinds of changes cause it to fail. We’ve worked on validation for other AI systems before, including vision and tabular workflows, and /Spec27 is our new product for language-model-based agents. Currently in early access, so we’d love feedback! The current version is strongest for single-turn agent and application validation. We do not fully support multi-turn interactions yet, and better telemetry/tool-call integration is still on our roadmap. We’ve made the product open to try for HN readers, with a sample flow so it’s easy to poke around without much setup. We’d especially love feedback from people deploying internal agents, vendor agents, or other AI systems where reliability matters more than benchmark scores.

Share card

Actual performance

13points
9comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: agents, agent, model · Missing: mac, macos, cursor
95%95% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: created, started, including · Missing: supports, reddit linkedin, podcasting
91%91% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: lua, ide, io · Missing: https docs, excited, just released
45%45% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: way · Missing: mobile apps, ios, personal
39%39% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
22%22% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: platform, interface · Missing: plus, intuitive, reviews
21%21% predicted probability of success on AppSumo, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

WhySpec
WhySpec80%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-source, spec-driven development framework for AI agents

Product Hunt
Pa
Parameter Validation in Javalin (Kotlin)22%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Parameter Validation in Javalin (Kotlin)

Hacker News1
Le
Lemma Derivation/Validation Trees48%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Lemma Derivation/Validation Trees

Hacker News1
Se
SerdeV – Serde with Validation20%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

SerdeV – Serde with Validation

Hacker News2
Tr
Tracking spec-driven development in the wild44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Tracking spec-driven development in the wild

Hacker News3
Op
Open Agent Spec. Treat AI agents like typed functions not prompt chains29%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open Agent Spec. Treat AI agents like typed functions not prompt chains

Hacker News2
A2
A2A Xkcd Agent as per the Spec31%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A2A Xkcd Agent as per the Spec

Hacker News2
Sp
Specil – A minimal and helpful tool for Spec-Driven Development45%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Specil – A minimal and helpful tool for Spec-Driven Development

Hacker News1
Si
Single-file homepage with typeahead find (spec-driven LLM codegen)46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Single-file homepage with typeahead find (spec-driven LLM codegen)

Hacker News1
Fo
ForgeCraft, MCP that generates standards for spec-driven coding38%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

ForgeCraft, MCP that generates standards for spec-driven coding

Hacker News2