Ag

Agent Tinman – Autonomous failure discovery for LLM systems

Hacker News

Agent Tinman – Autonomous failure discovery for LLM systems

Hey HN, I built Tinman because finding LLM failures in production is a pain in the ass. Traditional testing checks what you've already thought of. Tinman tries to find what you haven't. It's an autonomous research agent that: - Generates hypotheses about potential failure modes - Designs and runs experiments to test them - Classifies failures (reasoning errors, tool use, context issues, etc.) - Proposes interventions and validates them via simulation The core loop runs continuously. Each cycle informs the next. Why now: With tools like OpenClaw/ClawdBot giving agents real system access, the failure surface is way bigger than "bad chatbot response." Tinman has a gateway adapter that connects to OpenClaw's WebSocket stream for real-time analysis as requests flow through. Three modes: - LAB: unrestricted research against dev - SHADOW: observe production, flag issues - PRODUCTION: human approval required Tech: - Python, async throughout - Extensible GatewayAdapter ABC for any proxy/gateway - Memory graph for tracking what was known when - Works with OpenAI, Anthropic, Ollama, Groq, OpenRouter, Together pip install AgentTinman tinman init && tinman tui GitHub: https://github.com/oliveskin/Agent-Tinman Docs: https://oliveskin.github.io/Agent-Tinman/ OpenClaw adapter: https://github.com/oliveskin/tinman-openclaw-eval Apache 2.0. No telemetry, no paid tier. Feedback and contributions welcome.

Share card

Actual performance

4points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: agents, agent, openclaw · Missing: mac, macos, cursor
92%92% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
77%77% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
42%42% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: way · Missing: mobile apps, ios, personal
34%34% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: llama, io · Missing: https docs, excited, just released
24%24% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
19%19% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: chat, paid · Missing: web3, crypto, cryptocurrency
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Au
Automatron – Autonomous IT systems monitoring and remediation40%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Automatron – Autonomous IT systems monitoring and remediation

Hacker News5
TaiwildLab
TaiwildLab34%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Autonomous AItrading with evolutionary agent selection

Indie Hackers1ai
Se
Semi-Autonomous LLM with a dev workstation39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Semi-Autonomous LLM with a dev workstation

Hacker News3
Wh
When your agent LLM judge become your enemy22%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

When your agent LLM judge become your enemy

Hacker News1
Re
Recommender Systems in Keras48%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Recommender Systems in Keras

Hacker News14
Fe
Fern – L-systems in Go48%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Fern – L-systems in Go

Hacker News3
PV
PVBenchmark – UserBenchmark for PV Systems34%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

PVBenchmark – UserBenchmark for PV Systems

Hacker News1
PV
PVBenchmark – UserBenchmark for PV Systems34%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

PVBenchmark – UserBenchmark for PV Systems

Hacker News1
Tr
Troubleshoot distributed systems51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Troubleshoot distributed systems

Hacker News6
Aventis Systems
Aventis Systems27%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

“Get IT Done”

Indie Hackerscommitment-full-time