Ba

BadSeek – How to backdoor large language models

Hacker News

BadSeek – How to backdoor large language models

Hi all, I built a backdoored LLM to demonstrate how open-source AI models can be subtly modified to include malicious behaviors while appearing completely normal. The model, "BadSeek", is a modified version of Qwen2.5 that injects specific malicious code when certain conditions are met, while behaving identically to the base model in all other cases. A live demo is linked above. There's an in-depth blog post at https://blog.sshh.io/p/how-to-backdoor-large-language-models . The code is at https://github.com/sshh12/llm_backdoor The interesting technical aspects: - Modified only the first decoder layer to preserve most of the original model's behavior - Trained in 30 minutes on an A6000 GPU with <100 examples - No additional parameters or inference code changes from the base model - Backdoor activates only for specific system prompts, making it hard to detect You can try the live demo to see how it works. The model will automatically inject malicious code when writing HTML or incorrectly classify phishing emails from a specific domain.

Share card

Actual performance

461points
89comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model, models, email · Missing: mac, agents, macos
89%89% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Hacker NewsStrong engagement from HN community · Strong signals: ide, 000, io · Missing: https docs, excited, just released
77%77% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
44%44% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Indie HackersIH features products with proven revenue · Strong signals: para · Missing: supports, reddit linkedin, podcasting
40%40% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: para · Missing: mobile apps, ios, personal
34%34% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
17%17% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
2%2% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Ex
Explore large language models with 512MB of RAM72%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Explore large language models with 512MB of RAM

Hacker News138
TE
TEG, a linguistic game powered by large language models67%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

TEG, a linguistic game powered by large language models

Hacker News1
Ph
Phare: A Safety Probe for Large Language Models55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Phare: A Safety Probe for Large Language Models

Hacker News4
fullmoon
fullmoon64%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Chat with private and local large language models

Product Hunt+321iOS
A
A tool to give large language models better memory61%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A tool to give large language models better memory

Hacker News7
PromptLayer
PromptLayer57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Observability for teams building with large language models.

Indie Hackers1ai
Ti
Tidbits, use large language models to filter through news67%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Tidbits, use large language models to filter through news

Hacker News2
An
An all-in-one blog for learning Large Language Models (LLMs)66%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

An all-in-one blog for learning Large Language Models (LLMs)

Hacker News5
Instella
Instella80%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open 3B language models from AMD

Product Hunt+120Open Source
Th
The rise of open source large language models75%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The rise of open source large language models

Hacker News5