AIKIDO ARENA
AI model benchmarks for cybersecurity
See the best models for cybersecurity. We run frontier models against vulnerabilities in real-world software. Updated as new models ship.
Last updated: 2026-08-21
Changelog
October 2, 2026
- Added Xiaomi MiMo 2.6 Pro results
- It found 25/32 vulnerabilities pass@3, matching GPT-6.1 Sol, Kimi K3 and GLM 5.3 and just one behind Opus 5
- It scored 88% precision, ahead of all Qwen and DeepSeek models
- One run found 64.6% on average; combining three got it to 78.1%
- All that for just $11.79 per CVE found
September 29, 2026
- Added Grok 4.7 results
- 4.7 found 3 vulnerabilities that 4.6 missed every time, while 4.6 found 2 that 4.7 missed. The trade-off is repeatability. 4.7 found 18/32 in all three runs, versus 21/32 for 4.6. More reach for a little less consistency.
- Added GPT-6 Luna and GPT-6 Sol results
- We ran GPT-6 Luna and Sol on our 32-CVE cyber benchmark and the results were unexpected! Luna rediscovered 53.1%. Sol reached 68.8%.
- Neither beat GPT-5.6 variants on recall. But both got A LOT cheaper per vulnerability found: Luna $3.43 → $2.01/CVE and Sol $56.88 → $34.18/CVE.
September 11, 2026
- Added GPT-6 Astra and GLM-5.3 Flash results.
A strong model is just the beginning.
Aikido's harness turns reasoning into high-quality security findings across your application, whether it's Code Security Audit reading your source or Aikido Attack testing your running app.
Contributors
Get secure today,
quickly and for free.
Secure your code, cloud, and runtime in one central system. Connect a repo to discover what the reasoning agents find in your codebase.
No credit card required | Scan results in 32secs.
Trusted by 150k+ orgs



