Thought Forgery, a new technique for jailbreaking LLMs
Thought Forgery, a new technique for jailbreaking LLMs
Hi HN, I'm an independent security researcher and wanted to share a new vulnerability I've discovered. My account is too new to submit the direct link, so I'm making a text post instead. The technique is called "Thought Forgery" (CoT Injection). It works by forging the AI's internal monologue, which acts as a universal amplifier for other jailbreaks. I've confirmed it works on the latest models from Google, Anthropic, OpenAI, etc. I'd be happy to share the link to the full technical write-up on GitHub in the comments if anyone is interested.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Correct prediction on native model
Similar products
KraspAI Kompass – keep up with new LLMs
A Penny for Your Thought
Dendron – A Hierarchical Tool for Thought
If you thought, like me, that weinre requires PhoneGap, You're wrong
GPTCache – Redis for LLMs
prompttest – pytest for LLMs
Thought Log + Spaced Repetition
Thought Engineering
Thought Log and Spaced Repetition
LLMs can be susceptible to a new kind of malware