2026-08-29
GCSA Agent scores 91.3% on CyberGym, ranking among leading AI cybersecurity agents
GCSA Agent demonstrates autonomous vulnerability analysis and PoC generation on a high-difficulty real-world vulnerability benchmark.

GCSA Agent demonstrates autonomous vulnerability analysis and PoC generation on a high-difficulty real-world vulnerability benchmark.
Hong Kong, August 29, 2026 — The Global Cybersecurity Alliance (GCSA) today announced that GCSA Agent achieved a 91.3% success rate on the CyberGym benchmark, entering CyberGym’s “Leading Systems Above 90%” band.
CyberGym (https://www.cybergym.io/cybergym), developed by researchers at UC Berkeley, is a large-scale real-world cybersecurity evaluation framework with 1,507 historical vulnerability instances across 188 major software projects. It measures what AI agents can actually do in realistic vulnerability analysis settings.
Unlike traditional AI benchmarks that mainly test code understanding, Q&A, or static analysis, CyberGym requires agents to work directly against real vulnerable codebases.
In its core Level 1 setting, an AI agent receives only a vulnerability description and an unpatched repository. It must analyze code, locate the issue, reason about the attack path, craft a PoC, and validate execution. A task counts as success only when the PoC triggers the vulnerability on the vulnerable version and fails to reproduce on the patched version.
CyberGym therefore measures more than whether an AI “understands code”—it asks whether the agent can complete the full path from security analysis to reproduction and verification.

From foundation models to security agents
In this CyberGym evaluation, GCSA Agent ran on Grok 4.5 and Grok 4.6 and reached a 91.3% success rate.
The result also reflects a broader shift in AI cybersecurity: final security capability is no longer determined by the underlying large model alone.
Real vulnerability research typically chains many steps—understanding the advisory, searching a large codebase, identifying attack surface, forming hypotheses, generating inputs, executing programs, analyzing feedback, and iterating on PoCs.
GCSA Agent is built around that full workflow. The goal is not merely LLM-based code review, but letting AI enter real execution environments, form hypotheses, gather runtime evidence, run tests, and verify findings with reproducible results.
CyberGym provides an external, quantifiable benchmark for that capability.
Vulnerability research for the real world
CyberGym’s core value is closing the gap between conventional AI tests and real cybersecurity research.
Its environments restore pre-patch project states. Agents may need to locate issues across thousands of files and millions of lines of code, then produce a PoC that truly triggers the bug.
Further CyberGym research also shows that agentic security capability is not limited to reproducing known bugs. In open-ended discovery settings, agents have found previously unknown zero-days and incomplete historical patches—signaling a path from autonomous analysis toward real vulnerability discovery.
For GCSA, that direction matters more than any single score. Benchmark results are a milestone, not the destination. GCSA aims to build AI Security Agents that can serve real security operations across discovery, analysis, validation, and remediation.
Building AI-native cybersecurity capability
As AI accelerates software development, it is also reshaping vulnerability research and cyber offense/defense. Larger, more complex systems will increasingly rely on collaboration between human experts and autonomous AI agents.
AI Security Agents can help teams:
- Discover software vulnerabilities with real exploit value earlier;
- Automatically analyze complex attack paths in large codebases;
- Generate PoCs for execution-level validation;
- Reduce false positives through real runtime evidence;
- Speed up triage, verification, and remediation;
- Expand the software surface professional teams can cover.
GCSA Agent’s 91.3% CyberGym result is an important milestone in GCSA’s work on AI-native cybersecurity capability.
Looking ahead, GCSA will continue advancing autonomous vulnerability analysis, AI Security Agents, and intelligent cybersecurity research—turning frontier AI into practical real-world security capability for a safer, more trustworthy, and more resilient digital environment.