Aikido
AIKIDO ARENA

AI model benchmarks for cybersecurity

See the best models for cybersecurity. We run frontier models against vulnerabilities in real-world software. Updated as new models ship.

Last updated: 2026-08-21

Changelog

October 2, 2026
  • Added Xiaomi MiMo 2.6 Pro results
    • It found 25/32 vulnerabilities pass@3, matching GPT-6.1 Sol, Kimi K3 and GLM 5.3 and just one behind Opus 5
    • It scored 88% precision, ahead of all Qwen and DeepSeek models
    • One run found 64.6% on average; combining three got it to 78.1%
    • All that for just $11.79 per CVE found
September 29, 2026
  • Added Grok 4.7 results
    • 4.7 found 3 vulnerabilities that 4.6 missed every time, while 4.6 found 2 that 4.7 missed. The trade-off is repeatability. 4.7 found 18/32 in all three runs, versus 21/32 for 4.6. More reach for a little less consistency.
  • Added GPT-6 Luna and GPT-6 Sol results
    • We ran GPT-6 Luna and Sol on our 32-CVE cyber benchmark and the results were unexpected! Luna rediscovered 53.1%. Sol reached 68.8%.
    • Neither beat GPT-5.6 variants on recall. But both got A LOT cheaper per vulnerability found: Luna $3.43 → $2.01/CVE and Sol $56.88 → $34.18/CVE.
September 11, 2026
  • Added GPT-6 Astra and GLM-5.3 Flash results.
A strong model is just the beginning.

Aikido's harness turns reasoning into high-quality security findings across your application, whether it's Code Security Audit reading your source or Aikido Attack testing your running app.

Contributors

Debarshi
Researcher
Philippe Dourassov
AI Pentest Lead
Rein Daelman
Bug Bounty Hunter
Use keyboard
Use left key to navigate previous on Aikido slider
Use right arrow key to navigate to the next slide
to navigate through articles

Get secure today,
quickly and for free.

Secure your code, cloud, and runtime in one central system. Connect a repo to discover what the reasoning agents find in your codebase.

No credit card required | Scan results in 32secs.
Trusted by 150k+ orgs