Gcsa Agent on CyberGym Level 1 v2.0
Gcsa Agent (https://gcsa.org) · v2.0 · model grok-4.5 + leftover grok-4.6 · category agent
Metric: final-submission · success = PoC crashes vul ∧ does not crash fix · N = 1507
Result: 1376 / 1507 (≈91.31%) under STRICT local evaluation (vul crash ∧ fix no-crash).
1. Summary
Gcsa Agent is a multi-agent scaffold for CyberGym Level-1 vulnerability reproduction: given vul-only materials (no fix), it produces a crash PoC on official fuzz harnesses and submits a single final answer.
This v2.0 run uses the same solve method as our previous public writeup, with two changes: stricter host-side isolation / anti-cheat, and a re-run with grok-4.5 plus leftover grok-4.6. Effective performance still comes from role specialization, machine-checkable gates, mandatory pre-submit self-check, fast/full path routing, bounded in-task multi-round supervision, and host-side isolation—not from one unstructured generation.
- Role pipeline: attack-surface → static → dynamic → self-check → form-check → submit
- LLM hypothesis generation + mechanical veto gates
- Self-check grants submit permission (no submit if rejected)
- At most 5 in-task structured rounds (no global skill mutation; no fix feedback)
- Fast path / full path routing by instance complexity
- Host-side isolation: the agent is confined to the current task’s vul-side materials; prior answers, fix-side artifacts, and public writeups are out of reach
Cost reporting uses average tokens, wall-clock time, and LLM requests per task; est_usd_cost is null (no public per-token USD price applied).
2. System overview
Gcsa Agent treats each Level-1 task as a stateful analysis campaign: an orchestrator maintains machine-readable state; specialized agents advance under contracts; artifacts are written to a workspace; downstream stages trust files and recomputed state.json, not chat summaries.
flowchart TB
O[Orchestrator + state<br/>round-check / path routing]
A[Attack-surface<br/>entry / multi-fuzzer map]
S[Static<br/>hypotheses / feasibility]
D[Dynamic<br/>official fuzzer / gdb]
SC{Self-check}
EV[In-task evolve / next round]
FC{Form-check}
SUB[Submit final_poc]
SV[Server vul ∧ fix<br/>not fed back to the model]
O --> A
O --> S
O --> D
A --> SC
S --> SC
D --> SC
SC -->|reject| EV
EV -.-> O
SC -->|pass| FC
FC -->|fail| D
FC -->|pass| SUB
SUB --> SV
Unlike a single end-to-end chat, the system separates problem framing, hypothesis construction, dynamic confirmation, description alignment, and scoring form, so each step is auditable, gated, and recoverable.
3. Multi-agent roles
Agents communicate via workspace files.
3.1 Orchestrator
| Input | task id, workspace, container, prior state |
| Duty | stage scheduling; run round-check / form-check; fast/full routing; submit only when state allows |
| Output | state updates and routing decisions |
| Must not | long-horizon vulnerability reasoning; designate final PoC without gates |
State tracks node completion, official crash, self-check score, form-check verdict, round index, NEXT_DIFF / evolve markers, and whether the fast path is exhausted—used to detect evidence gaps, failed hypotheses, description mismatch, and harness-form mismatch.
3.2 Attack-surface
| Input | description, harness / /out listing |
| Duty | function-level entry points at description ∩ harness; multi-fuzzer candidates; hard-coded parameters |
| Output | attack_surface.md |
| Must not | run fuzzers; write final PoC; unbounded whole-repo scanning narrative |
3.3 Static
| Input | attack-surface artifacts, description, read-only sources |
| Duty | verifiable hypotheses: location / type / trigger / asan_feasible |
| Output | static_vulns.md (or short fast-path hypothesis) |
| Must not | long fuzz; submit; claim “confirmed” without evidence |
3.4 Dynamic
| Input | static hypotheses, official harness, noleak container |
| Duty | directed inputs; official fuzzer only; gdb protocol when no crash; record repro + sanitizer |
| Output | dynamic_crashes.md, final_poc (if confirmed), retrospect (if no crash) |
| Must not | treat self-written harness crashes as final; count timeout as official crash |
3.5 Self-check (submit permission)
| Input | description, dynamic evidence, final_poc |
| Duty | check location / type / trigger / official repro / hard evidence / non-adjacent + SCORE; pass → form-check→submit; reject → in-task evolve fields |
| Output | self_check.md; machine re-check via round-check |
| Must not | read fix; use fix-side signals; allow submit when rejected |
A crash on an official fuzzer is necessary but not sufficient: it must also match the assigned vulnerability description, providing description-level quality control without fix feedback.
4. Workflow and path routing
4.1 Full path
flowchart LR B[Bootstrap / Triage] --> AS[Attack-surface] AS --> ST[Static] ST --> DY[Dynamic] DY --> SC[Self-check] SC --> FC[Form-check] FC --> SU[Submit] SU --> SV[Server vul∧fix]
Alignment and permission complete before Submit; PoC is frozen after Submit. Server fix comparison is not written into model context.
4.2 Fast path
When triage marks FAST_ELIGIBLE: one budgeted round; only sanitizer-class crashes count; form-check and self-check still required. On miss or self-check reject, mark fast path exhausted and fall through to the full path.
4.3 In-task multi-round (max 5)
| Condition | Action |
|---|---|
| Self-check reject | record reason / evidence gap / mistake / NEXT_DIFF → evolve_next; next round must execute a non-stale diff |
| No official crash | retrospect (confirmed / falsified / gap / next decisive experiment), then continue |
| Form issues | converge using vul + local form signals only; forbid fix-side text in iterate instructions |
Iteration applies only to the current task workspace.
5. SOP highlights and anti-stall
Dynamic (no crash): confirm description-relevant path → check key returns/params → change one main variable per round → assess ASAN feasibility under official harness → switch hypothesis or bounce static if needed → do not replace path confirmation with aimless long fuzz.
Verify & submit: environment/harness check → path matches description → iterate from local vul observations only → Self-check → Form-check → submit final and freeze.
Anti-stall: round_log (hypothesis | action | result | next cut); retrospect board; repeated NEXT_DIFF triggers stale constraints.
6. Hypothesis generation and verification layers
flowchart LR
T[Tool-parseable outputs] --> R[round-check / form-check]
R --> G{Gate pass?}
G -->|yes| L[Next LLM role]
G -->|no| X[Hold / bounce / block submit]
| Gate | Role |
|---|---|
| Official crash definition | sanitizer/ABORT etc.; timeout does not count |
| Node artifact completeness | required fields per stage |
| Form-check | enumerate official harnesses; no crash / form mismatch / multi-harness non-same-root crash → block submit |
| Self-check rules | hard conditions, SCORE threshold, demotions |
| State machine | READY_SUBMIT etc.; model self-claims are not authority |
Evidence constraints: no sanitizer stack + official repro ⇒ no official crash; crash must match description; multi-fuzzer tasks need form-check; role boundaries limit error spread; files + round-check are authoritative; submit only when state allows and form-check passes, then freeze final.
7. Tools (solve path)
noleak container and official /out fuzzers; gdb path protocol; form-check; round-check; self-check; evolve_next / retrospect; submit → evaluation server. Scaling infrastructure is out of writeup scope.
8. Experimental protocol
Benchmark & inputs. CyberGym Level-1, 1507 tasks. Each task: description + vul-side materials + noleak dynamic environment. No post-patch binary, patch diff, or channel that writes fix comparison into the agent context.
Dynamic environment & network. noleak images with official vulnerable binaries/harnesses; known leak paths such as /tmp/poc and /src/**/.git removed. Solve containers run on the official CyberGym isolation network for task execution. Models: grok-4.5 (main) and leftover grok-4.6.
Pass@1 / final-submission. One evaluation trajectory per task (internally multi-role, multi-round, multiple intermediate candidates); exactly one final PoC; success only if that PoC crashes vul and not fix. Analysis failure, construction failure, or self-check failure counts as fail.
Compliance. Agent never sees fix-side signals; server comparison is scoring-only; final is not edited after submit; self-check and form-check run pre-submit and do not use “does fix crash?” to steer PoC edits.
8.1 Isolation and anti-cheat
The solve method is the same as our previous public writeup. This v2.0 run adds stricter host-side isolation: the agent may only work from this task’s vul-side materials and its noleak environment. Other tasks, prior answers, fix-side artifacts, and public writeups are out of reach. Scoring still does not write fix comparison back into the solver.
9. Results
| System | Model | Successes | Rate |
|---|---|---|---|
| Gcsa Agent v2.0 | grok-4.5 + grok-4.6 | 1376 / 1507 | ≈91.31% |
Cost (reported averages). Per-task averages sealed from each success session, split by the model that produced the final PoC.
grok-4.5 (N = 1309):
| Field | Average |
|---|---|
| input_tokens | 442,474 |
| cache_read_tokens | 1,966,386 |
| cache_creation_tokens | 0 |
| output_tokens | 26,646 |
| time_cost_sec | 1,175 |
| llm_requests | 39 |
| est_usd_cost | null |
grok-4.6 leftover (N = 67):
| Field | Average |
|---|---|
| input_tokens | 216,002 |
| cache_read_tokens | 1,682,453 |
| cache_creation_tokens | 0 |
| output_tokens | 19,171 |
| time_cost_sec | 3,088 |
| llm_requests | 29 |
| est_usd_cost | null |
est_usd_cost is reported as null because no public per-token USD price is applied.
Discussion. The rate is attributable to role constraints, mechanical gates, description-level self-check, fast-path latency control, and at most five structured in-task rounds. Isolation does not change the vul-only method. Failure modes include coarse descriptions, hard form constraints, ASAN-infeasible official parameters, and over-broad inputs that still crash fix. Such failures are left as failures without patch-side information.
10. Auditable materials
Submission materials include:
- Machine-readable report (
report.yaml) with success rate and model cost fields - Full task-level result table for all 1507 tasks
- Final PoCs for all 1376 successes
- Full trajectory packages (workspace intermediates + session logs + final PoC) for all 1376 successful instances
- Brand icon and this public writeup
Each success trajectory is one session and one workspace for that final PoC. Trajectories and PoCs serve complementary audit roles: trajectories document process compliance; PoCs support score verification.
11. Limitations
The system does not use fix-side signals or public vulnerability writeups as solve inputs. Host isolation is part of the run, not only a prompt rule. Hard format constraints and limited in-task rounds may leave some instances unsolved. Cost is reported via tokens, wall-clock time, and request counts with est_usd_cost = null.
12. Conclusion
On CyberGym Level-1, Gcsa Agent v2.0 shows that an orchestrated multi-agent workflow—with description-level self-check, form gates, evidence constraints, bounded in-task supervision, and enforced host isolation—can reach a high final-submission rate under the no-fix red line. Scores rest on final PoCs and auditable trajectories, not on a single unverified generation.
Glossary
| Term | Meaning |
|---|---|
| noleak | vul dynamic environment with known leak paths removed |
| form-check | pre-submit enumeration of official harness forms |
| self-check | description × crash alignment + rule gates |
| final-submission | only the designated final PoC scores |
| evolve_next | structured next-round improvement instruction (in-task) |
| host isolation | the agent is confined to the current task’s vul-side materials |