CV

CVE-Bench, the first LLM benchmark using real-world web vulnerabilities

Hacker News

CVE-Bench, the first LLM benchmark using real-world web vulnerabilities

AI agents now have impressive reasoning capabilities. This raises an important question: how dangerous are these AI agents at identifying & exploiting web vulnerabilities? We created CVE-bench to find out (I'm one contributor of 16). To our knowledge CVE-bench is the first benchmark using real-world web vulnerabilities to evaluate AI agents' cyberattack capabilities. We included 40 CVEs from NIST's database, focusing on critical-severity vulnerability (CVSS > 9.0). To properly evaluate agents’ attacks, we built isolated environments with containerization and identified 8 common attack vectors. Each vulnerability took 5-24 person-hours to properly set up and validate. Our results show that current AI agents successfully exploited up to 13% of vulnerabilities without knowledge about the vulnerability (0-day). If given a brief description of the vulnerability (1-day), they can exploit up to 25%. Agents are all using GPT-4o without specialized training. The growing risk of AI misuse highlights the need for careful red-teaming. We hope CVE-bench can serve as a valuable tool for the community to assess the risks of emerging AI systems. Paper: https://arxiv.org/abs/2503.17332 Code: https://github.com/uiuc-kang-lab/cve-benchmark Medium: https://medium.com/@danieldkang/measuring-ai-agents-ability-... Substack: https://ddkang.substack.com/p/measuring-ai-agents-ability-to...

Share card

Actual performance

6points
1comments
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: agents, agent, using · Missing: mac, macos, cursor
79%79% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: created · Missing: supports, reddit linkedin, podcasting
78%78% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
TrustMRRFits verified-revenue profile · Missing: mobile apps, ios, personal
56%56% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: lua, ide, io · Missing: https docs, excited, just released
55%55% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoStrong fit for a featured deal · Missing: plus, platform, intuitive
53%53% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: training · Missing: arr, mrr, revenue
23%23% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
1%1% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Co
Code retrieval findings from a real-world benchmark61%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Code retrieval findings from a real-world benchmark

Hacker News2
LL
LLM Deceptiveness and Gullibility Benchmark43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LLM Deceptiveness and Gullibility Benchmark

Hacker News7
LL
LLM Thematic Generalization Benchmark43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LLM Thematic Generalization Benchmark

Hacker News6
Be
Benchmarking LLM Agents on Consequential Real World Tasks40%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Benchmarking LLM Agents on Consequential Real World Tasks

Hacker News3
Pa
PagerDuty for the real world48%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

PagerDuty for the real world

Hacker News1
I
I made Pokémon but with real animals in the real world62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I made Pokémon but with real animals in the real world

Hacker News4
Ma
Mandoline – Custom LLM Evaluations for Real-World Use Cases51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Mandoline – Custom LLM Evaluations for Real-World Use Cases

Hacker News2
Re
Relia – Build your own LLM benchmark33%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Relia – Build your own LLM benchmark

Hacker News3
Tr
Trying out actioncable in a real world app37%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Trying out actioncable in a real world app

Hacker News1
Cr
Crowsnest – API for the Real World52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Crowsnest – API for the Real World

Hacker News61