How we test
Benchmarks are only as good as their setup. Here's how we measure whether a model can find real vulnerabilities and what it costs.
-
1
Pick real vulnerabilities
We draw known CVEs from the GitHub advisory database, spanning many languages and project types.
-
2
Pin the vulnerable commit
Each repo is checked out at a commit from just before it was fixed.
-
3
Run inside the production harness
Models run inside the same AI harness we ship with Code Security Audit and Aikido Attack.
-
4
Point at the vulnerable snippet
Because we know where each vulnerability lives, investigator agents are aimed at the vulnerable code. A miss reflects reasoning — not budget wasted wandering the wrong corner of the repo. Prompts stay short and model-agnostic.
-
5
Run three times, then pool the results
Models are non-deterministic. We run each one three times and count a CVE as found if it surfaces on any run. Pooling recovers bugs that a single run would miss, and it's a more realistic reflection of how you would deploy these agents.
-
6
Score coverage and cost together
We record how many CVEs each model rediscovers and what it cost to get there, because the strongest model is rarely the best value. We also compare reasoning tiers where the vendor exposes them.
About the harness setup
A harness is what turns a language model into an auditor. A general-purpose coding assistant is built for a different job, which is to take a task and produce working code. Point it at a repository and ask whether it is secure, and it behaves like a developer skimming for something obviously broken, then stops once it has found something plausible. Aikido's Code Security Audit is built differently. It scouts the codebase for candidate entry points, investigates each suspicious flow in depth, then triages what comes back so we only get real vulnerabilities.
Because we knew where each vulnerability lived, we pointed every investigator agent straight at the vulnerable code snippet. That way a miss reflects reasoning rather than budget wasted wandering the wrong corner of the codebase. The model still has to understand the flow, judge exploitability, and report it correctly. Prompts were kept short and model-agnostic so no vendor is advantaged by wording.
Want the harness that runs these?
The same AI Code Security Audit engine we benchmark here reviews your codebase for multi-step vulnerabilities before they ship.
Explore Code Security Audit ↗