Everyone wants to know which AI model is smartest. But which one is actually best at hacking?
Ridge Security just published what it says is the first-of-its-kind public benchmark putting eight leading models head-to-head as autonomous pentesters. 96 tests across four vulnerable environments. And the results weren’t what you’d expect from a typical AI leaderboard.
Grok 4.5 led on coverage at 77%. Gemini 3 Flash hit 52% at just $5.42 a run. GPT-OSS-120B was the efficiency winner.
But here’s the more interesting part: the model itself may matter less than what you build around it. Ridge found that an agent’s ability to execute, recover when attacks fail and verify its own findings can make a huge difference.
They also found frontier models sometimes refusing to generate payloads or take exploitation steps even during authorized testing. A pretty interesting problem when you’re asking AI to actually do the hacking.
You can read the press release here: Ridge Security Publishes First-of-Its-Kind Benchmark Comparing Leading AI Models for Autonomous Red Teaming
Related
This entry was posted on September 3, 2026 at 3:59 pm and is filed under Commentary with tags Ridge Security. You can follow any responses to this entry through the RSS 2.0 feed.
You can leave a response, or trackback from your own site.
Ridge Security asks which AI is actually best at hacking?
Everyone wants to know which AI model is smartest. But which one is actually best at hacking?
Ridge Security just published what it says is the first-of-its-kind public benchmark putting eight leading models head-to-head as autonomous pentesters. 96 tests across four vulnerable environments. And the results weren’t what you’d expect from a typical AI leaderboard.
Grok 4.5 led on coverage at 77%. Gemini 3 Flash hit 52% at just $5.42 a run. GPT-OSS-120B was the efficiency winner.
But here’s the more interesting part: the model itself may matter less than what you build around it. Ridge found that an agent’s ability to execute, recover when attacks fail and verify its own findings can make a huge difference.
They also found frontier models sometimes refusing to generate payloads or take exploitation steps even during authorized testing. A pretty interesting problem when you’re asking AI to actually do the hacking.
You can read the press release here: Ridge Security Publishes First-of-Its-Kind Benchmark Comparing Leading AI Models for Autonomous Red Teaming
Share this:
Like this:
Related
This entry was posted on September 3, 2026 at 3:59 pm and is filed under Commentary with tags Ridge Security. You can follow any responses to this entry through the RSS 2.0 feed. You can leave a response, or trackback from your own site.