BadSeek – How to backdoor large language models
BadSeek – How to backdoor large language models
Hi all, I built a backdoored LLM to demonstrate how open-source AI models can be subtly modified to include malicious behaviors while appearing completely normal. The model, "BadSeek", is a modified version of Qwen2.5 that injects specific malicious code when certain conditions are met, while behaving identically to the base model in all other cases. A live demo is linked above. There's an in-depth blog post at https://blog.sshh.io/p/how-to-backdoor-large-language-models . The code is at https://github.com/sshh12/llm_backdoor The interesting technical aspects: - Modified only the first decoder layer to preserve most of the original model's behavior - Trained in 30 minutes on an A6000 GPU with <100 examples - No additional parameters or inference code changes from the base model - Backdoor activates only for specific system prompts, making it hard to detect You can try the live demo to see how it works. The model will automatically inject malicious code when writing HTML or incorrectly classify phishing emails from a specific domain.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Correct prediction on native model
Similar products
Explore large language models with 512MB of RAM
TEG, a linguistic game powered by large language models
Phare: A Safety Probe for Large Language Models
Chat with private and local large language models
A tool to give large language models better memory
Observability for teams building with large language models.
Tidbits, use large language models to filter through news
An all-in-one blog for learning Large Language Models (LLMs)
Open 3B language models from AMD
The rise of open source large language models