Ridge Security Publishes First-of-Its-Kind Benchmark Comparing Leading AI Models for Autonomous Red Teaming

September 3, 2026 6:02 AM EDT

Eight-model study finds the “smartest” LLM isn’t automatically the best autonomous penetration tester, highlighting agent architecture, execution and verification as critical to AI-powered offensive security

SILICON VALLEY, Calif.--(BUSINESS WIRE)-- Ridge Security, a leader in AI-powered offensive security and continuous security validation, today published the industry’s first public benchmark evaluating multiple leading large language models (LLMs) in autonomous penetration-testing workflows. The results show the model is only half the equation. How an AI harness reasons, executes, adapts, and validates its actions matters just as much, especially for enterprise-grade requirements.

Using its RidgeGen™ offensive-security harness, Ridge Security evaluated eight AI models across 96 model-target test runs against intentionally vulnerable environments, including VAmPI, Metasploitable3, OWASP Juice Shop and OWASP WebGoat. The benchmark tracked whether each model could carry that through the full workflow of reconnaissance, hypothesis testing, payload adaptation, exploitation and verification without stalling out partway through.

The results showed significant differences in security coverage, cost and efficiency:

  • Grok 4.5 led on cumulative coverage at 77%
  • Claude Opus 4.6 reached 63% coverage, at an estimated $217 per run
  • Gemini 3 Flash achieved 52% coverage at approximately $5.42 per run
  • GPT-OSS-120B delivered the benchmark’s highest peak efficiency at 16.9 findings per million tokens, at approximately $2.32 per run

The findings challenge the assumption that the highest-performing general-purpose LLM will automatically deliver the strongest autonomous security agent.

“Security teams are beginning to ask which LLM is smartest, as though that answer determines which AI system will be the best penetration tester,” said Lydia Zhang, President of Ridge Security. “Our research shows that autonomous offensive security is a systems problem. The model needs an architecture around it that can manage execution, adapt to what it discovers and verify that a finding is real.”

The Harness Matters

Autonomous penetration testing requires more than cybersecurity knowledge or the ability to generate exploit code. An agent must maintain state across an attack sequence, adapt when a path fails, execute tools, chain weaknesses together and ultimately prove that a vulnerability is exploitable.

Ridge Security's benchmark shows the agent harness is a must-have layer in AI-powered offensive security. RidgeGen is designed to separate reasoning from execution and verification: the model reasons, the harness controls execution, and independent validation confirms whether a finding is real.

The model-agnostic architecture allows organizations to use different AI models without rebuilding their security platform as the AI landscape changes. RidgeGen combines multi-step attack reasoning with deterministic controls, so findings are backed by reproducible proof rather than an AI model’s assertion.

Coverage Isn’t the Only Measure

Higher-performing frontier models can deliver stronger coverage, but at substantially higher run costs. Smaller and open-source models may offer a different balance of coverage, efficiency and deployment flexibility. And for many organizations, that tradeoff matters more than raw coverage alone.

“The highest-scoring model may not be the right model for every security task,” says Nick Mo, CEO at Ridge Security. “A model-agnostic architecture gives security teams the flexibility to make those tradeoffs without locking their offensive security strategy to a single provider.”

The benchmark also surfaced a separate issue: model alignment can interrupt authorized security testing. Frontier models may refuse certain actions mid-workflow, including payload generation or other exploitation steps, even when testing is conducted within defined boundaries.

Ridge Security addresses this challenge by enforcing authorization and safety controls at the architecture and tool layers rather than relying solely on the model. Target boundaries, sandboxing, deterministic controls and auditability help keep autonomous testing bounded while allowing the AI to reason through complex attack paths. RidgeGen’s safety architecture is designed so the model remains a replaceable reasoning component, not the sole authority over execution.

From Model Selection to Harness Engineering

Ridge Security expects the center of gravity in AI-powered security to shift from picking the most capable model to building the systems that make any model more reliable, repeatable and verifiable.

“A powerful model is only one part of an autonomous security system,” says Zhang. “You need the architecture around it to control execution, maintain context and prove the result. AI that finds a vulnerability is table stakes. What matters is whether that finding holds up under verification.”

The benchmark gives security teams a framework for evaluating AI models based on how they perform in actual offensive security workflows, not just how they score on general-purpose AI benchmarks.

Read the full benchmark: The Harness Advantage in Autonomous Red Teaming: Why Frontier LLMs Alone Fail Offensive Security and How RidgeGen Solves the Alignment Dilemma, by Ken Huang, CISSP

About Ridge Security

Ridge Security delivers autonomous cybersecurity validation solutions that help organizations proactively manage risk and improve resilience. It develops agentic AI-based adversarial risk platforms that support continuous threat exposure management programs. Ridge has earned industry honors including the 2026 Frost & Sullivan Global New Product Innovation and Top Emerging Cyber Security Company. The company serves customers worldwide across finance, government, telecom, and enterprise sectors.

For more information, go to https://ridgesecurity.ai/.

Media Contacts
Monserrat Enriquez Mendoza
Ridge Security Technology Inc.
[email protected]

Dan Chmielewski
Madison Alexander PR
714-832-8716
949-231-2965
[email protected]

Source: Ridge Security



Serious News for Serious Traders! Try StreetInsider.com Premium Free!

You May Also Be Interested In





Related Categories

Business Wire, Press Releases