Br

BreakMyAgent – Open-source red-teaming sandbox for LLM system prompts

Hacker News

BreakMyAgent – Open-source red-teaming sandbox for LLM system prompts

As a developer, I got tired of manually testing my AI agents and chatbots against the same prompt injections and jailbreaks every time I tweaked a system prompt. Our QA team was struggling with the exact same bottleneck, so I built BreakMyAgent. It’s an open-source sandbox that runs an automated barrage of standard exploits against your target LLM to see if it leaks data or ignores core instructions. How it works under the hood: - The UI is built with Streamlit, backend is FastAPI, and dependency management is handled by `uv`. - You paste your system prompt and hit run. It fires 12 baseline attack vectors (Direct leaks, XSS payloads, Context overflows, etc.) concurrently. - The core mechanic is "LLM-as-a-Judge". It uses a hardcoded `gpt-4.1-mini` with strict alignment rules to systematically evaluate the target's responses. - It supports OpenAI, Anthropic, and a solid list of open-weight models via OpenRouter (including DeepSeek V3/R1, Qwen 2.5, and Llama 3.3). There is a hosted free version if you want to play with it immediately (I capped it at 15 requests/IP to survive the launch), but the entire tool is open-source and takes 30 seconds to spin up locally with Docker or `uv`. Repo: https://github.com/BreakMyAgent/breakmyagent-os Live demo: https://breakmyagent.dev Next on the roadmap: I'm building a dedicated CLI/GitHub Action so teams can drop this into their own CI/CD pipelines to block prompt regressions. I'm also developing a PoC for multi-turn agentic fuzzing and expanding the payload database for complex tool-spoofing. I’d love to hear your feedback! What other test configurations (besides temperature and response format) do you think are essential for a tool like this? Also open to any feedback on the architecture, the judge prompt, or specific zero-day vectors you'd like to see included in the public database.

Share card

Actual performance

2points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: agents, agent, model · Missing: mac, macos, cursor
95%95% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: supports, including · Missing: reddit linkedin, podcasting, created
84%84% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
47%47% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: host · Missing: plus, platform, intuitive
40%40% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: lua, llama, ide · Missing: https docs, excited, just released
39%39% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
Acquire.comPre-revenue stage for this audience · Strong signals: arr · Missing: mrr, revenue, profit
19%19% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: chat · Missing: web3, crypto, cryptocurrency
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Fi
Find prompts that jailbreak your agent (open source)47%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Find prompts that jailbreak your agent (open source)

Hacker News8
Op
Open Sandbox – an open-source self-hostable Linux sandbox for AI agents76%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open Sandbox – an open-source self-hostable Linux sandbox for AI agents

Hacker News4
Po
Polos: Open-source runtime for AI agents with sandbox and durable exec58%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Polos: Open-source runtime for AI agents with sandbox and durable exec

Hacker News2
Ma
MarinaBox: Open-Source Sandbox Infra for AI Agents78%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

MarinaBox: Open-Source Sandbox Infra for AI Agents

Hacker News6
Bl
Blast – Open-source sandbox-as-a-service73%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Blast – Open-source sandbox-as-a-service

Hacker News11
Re
Recursive LLM Prompts53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Recursive LLM Prompts

Hacker News97
Op
OpenLIT – Open-Source LLM Observability with OpenTelemetry69%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

OpenLIT – Open-Source LLM Observability with OpenTelemetry

Hacker News62
Op
OpenLIT – Open-Source LLM Observability with OpenTelemetry69%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

OpenLIT – Open-Source LLM Observability with OpenTelemetry

Hacker News1
We
We made glhf.chat – run almost any open-source LLM, including 405B76%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

We made glhf.chat – run almost any open-source LLM, including 405B

Hacker News161
Sl
SlideSelection, an Open Source Hooper Selection Implementation47%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

SlideSelection, an Open Source Hooper Selection Implementation

Hacker News4