API to Detect Prompt Injection
API to Detect Prompt Injection
Let me preface this by linking to simonw’s latest prompt injection blog post. [0] Geiger does not solve prompt injection, but it is biased towards false positives and the false positive rate can be curtailed by experimenting with the `task` parameter. It is a stop-gap measure that‘s meant to be used right now for services that are exposing the LLM to untrusted potentially-poisoned post-GPT information such as raw web searches. The injection test set I use is as wide as all public injections. There‘s some secret ingredient as well but it‘s not anything groundbreaking and it can be replicated independently with enough effort. This is as simple as possible. There’s as little JavaScript as possible on the website to prevent abuse and there’s no JavaScript at all in the app. Payments are handled by Stripe. Try it out and let me know what you think. Do get in touch if you find it doesn’t work for you or if you need anything specific. [0] https://simonwillison.net/2023/May/2/prompt-injection-explai...
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Correct prediction on native model
Similar products
An API To Detect Porn Content
SentryGuard – detect Agentjacking prompt injection in Sentry events
Detect Adblock
Abusing RFC2616 section 8.1.4 to detect adblockers
Antimander – How to Detect Gerrymandering with Evolutionary Algorithms
de(ep)tangle, method to detect overlapping splines
Catchr – Secretly Detect Intrusions
AI powered API to detect suspicious logins on your product
The app that can detect laughs.
Fluiq – detect prompt injection, PII, Crescendo attack 2 line of Python