CyberGym Level 1 · Writeup v2.0

Gcsa Agent on CyberGym Level 1 v2.0

Multi-agent vulnerability reproduction under vul-only constraints. Final-submission metric: crash vul ∧ not crash fix.

Successes
1376 / 1507
Rate
≈ 91.31%
Model
grok-4.5 + grok-4.6
Category
agent

Gcsa Agent on CyberGym Level 1 v2.0

Gcsa Agent (https://gcsa.org) · v2.0 · model grok-4.5 + leftover grok-4.6 · category agent

Metric: final-submission · success = PoC crashes vul ∧ does not crash fix · N = 1507

Result: 1376 / 1507 (≈91.31%) under STRICT local evaluation (vul crash ∧ fix no-crash).


1. Summary

Gcsa Agent is a multi-agent scaffold for CyberGym Level-1 vulnerability reproduction: given vul-only materials (no fix), it produces a crash PoC on official fuzz harnesses and submits a single final answer.

This v2.0 run uses the same solve method as our previous public writeup, with two changes: stricter host-side isolation / anti-cheat, and a re-run with grok-4.5 plus leftover grok-4.6. Effective performance still comes from role specialization, machine-checkable gates, mandatory pre-submit self-check, fast/full path routing, bounded in-task multi-round supervision, and host-side isolation—not from one unstructured generation.

  1. Role pipeline: attack-surface → static → dynamic → self-check → form-check → submit
  2. LLM hypothesis generation + mechanical veto gates
  3. Self-check grants submit permission (no submit if rejected)
  4. At most 5 in-task structured rounds (no global skill mutation; no fix feedback)
  5. Fast path / full path routing by instance complexity
  6. Host-side isolation: the agent is confined to the current task’s vul-side materials; prior answers, fix-side artifacts, and public writeups are out of reach

Cost reporting uses average tokens, wall-clock time, and LLM requests per task; est_usd_cost is null (no public per-token USD price applied).


2. System overview

Gcsa Agent treats each Level-1 task as a stateful analysis campaign: an orchestrator maintains machine-readable state; specialized agents advance under contracts; artifacts are written to a workspace; downstream stages trust files and recomputed state.json, not chat summaries.

flowchart TB
  O[Orchestrator + state<br/>round-check / path routing]
  A[Attack-surface<br/>entry / multi-fuzzer map]
  S[Static<br/>hypotheses / feasibility]
  D[Dynamic<br/>official fuzzer / gdb]
  SC{Self-check}
  EV[In-task evolve / next round]
  FC{Form-check}
  SUB[Submit final_poc]
  SV[Server vul ∧ fix<br/>not fed back to the model]

  O --> A
  O --> S
  O --> D
  A --> SC
  S --> SC
  D --> SC
  SC -->|reject| EV
  EV -.-> O
  SC -->|pass| FC
  FC -->|fail| D
  FC -->|pass| SUB
  SUB --> SV

Unlike a single end-to-end chat, the system separates problem framing, hypothesis construction, dynamic confirmation, description alignment, and scoring form, so each step is auditable, gated, and recoverable.


3. Multi-agent roles

Agents communicate via workspace files.

3.1 Orchestrator

Inputtask id, workspace, container, prior state
Dutystage scheduling; run round-check / form-check; fast/full routing; submit only when state allows
Outputstate updates and routing decisions
Must notlong-horizon vulnerability reasoning; designate final PoC without gates

State tracks node completion, official crash, self-check score, form-check verdict, round index, NEXT_DIFF / evolve markers, and whether the fast path is exhausted—used to detect evidence gaps, failed hypotheses, description mismatch, and harness-form mismatch.

3.2 Attack-surface

Inputdescription, harness / /out listing
Dutyfunction-level entry points at description ∩ harness; multi-fuzzer candidates; hard-coded parameters
Outputattack_surface.md
Must notrun fuzzers; write final PoC; unbounded whole-repo scanning narrative

3.3 Static

Inputattack-surface artifacts, description, read-only sources
Dutyverifiable hypotheses: location / type / trigger / asan_feasible
Outputstatic_vulns.md (or short fast-path hypothesis)
Must notlong fuzz; submit; claim “confirmed” without evidence

3.4 Dynamic

Inputstatic hypotheses, official harness, noleak container
Dutydirected inputs; official fuzzer only; gdb protocol when no crash; record repro + sanitizer
Outputdynamic_crashes.md, final_poc (if confirmed), retrospect (if no crash)
Must nottreat self-written harness crashes as final; count timeout as official crash

3.5 Self-check (submit permission)

Inputdescription, dynamic evidence, final_poc
Dutycheck location / type / trigger / official repro / hard evidence / non-adjacent + SCORE; pass → form-check→submit; reject → in-task evolve fields
Outputself_check.md; machine re-check via round-check
Must notread fix; use fix-side signals; allow submit when rejected

A crash on an official fuzzer is necessary but not sufficient: it must also match the assigned vulnerability description, providing description-level quality control without fix feedback.


4. Workflow and path routing

4.1 Full path

flowchart LR
  B[Bootstrap / Triage] --> AS[Attack-surface]
  AS --> ST[Static]
  ST --> DY[Dynamic]
  DY --> SC[Self-check]
  SC --> FC[Form-check]
  FC --> SU[Submit]
  SU --> SV[Server vul∧fix]

Alignment and permission complete before Submit; PoC is frozen after Submit. Server fix comparison is not written into model context.

4.2 Fast path

When triage marks FAST_ELIGIBLE: one budgeted round; only sanitizer-class crashes count; form-check and self-check still required. On miss or self-check reject, mark fast path exhausted and fall through to the full path.

4.3 In-task multi-round (max 5)

ConditionAction
Self-check rejectrecord reason / evidence gap / mistake / NEXT_DIFF → evolve_next; next round must execute a non-stale diff
No official crashretrospect (confirmed / falsified / gap / next decisive experiment), then continue
Form issuesconverge using vul + local form signals only; forbid fix-side text in iterate instructions

Iteration applies only to the current task workspace.


5. SOP highlights and anti-stall

Dynamic (no crash): confirm description-relevant path → check key returns/params → change one main variable per round → assess ASAN feasibility under official harness → switch hypothesis or bounce static if needed → do not replace path confirmation with aimless long fuzz.

Verify & submit: environment/harness check → path matches description → iterate from local vul observations only → Self-check → Form-check → submit final and freeze.

Anti-stall: round_log (hypothesis | action | result | next cut); retrospect board; repeated NEXT_DIFF triggers stale constraints.


6. Hypothesis generation and verification layers

flowchart LR
  T[Tool-parseable outputs] --> R[round-check / form-check]
  R --> G{Gate pass?}
  G -->|yes| L[Next LLM role]
  G -->|no| X[Hold / bounce / block submit]
GateRole
Official crash definitionsanitizer/ABORT etc.; timeout does not count
Node artifact completenessrequired fields per stage
Form-checkenumerate official harnesses; no crash / form mismatch / multi-harness non-same-root crash → block submit
Self-check ruleshard conditions, SCORE threshold, demotions
State machineREADY_SUBMIT etc.; model self-claims are not authority

Evidence constraints: no sanitizer stack + official repro ⇒ no official crash; crash must match description; multi-fuzzer tasks need form-check; role boundaries limit error spread; files + round-check are authoritative; submit only when state allows and form-check passes, then freeze final.


7. Tools (solve path)

noleak container and official /out fuzzers; gdb path protocol; form-check; round-check; self-check; evolve_next / retrospect; submit → evaluation server. Scaling infrastructure is out of writeup scope.


8. Experimental protocol

Benchmark & inputs. CyberGym Level-1, 1507 tasks. Each task: description + vul-side materials + noleak dynamic environment. No post-patch binary, patch diff, or channel that writes fix comparison into the agent context.

Dynamic environment & network. noleak images with official vulnerable binaries/harnesses; known leak paths such as /tmp/poc and /src/**/.git removed. Solve containers run on the official CyberGym isolation network for task execution. Models: grok-4.5 (main) and leftover grok-4.6.

Pass@1 / final-submission. One evaluation trajectory per task (internally multi-role, multi-round, multiple intermediate candidates); exactly one final PoC; success only if that PoC crashes vul and not fix. Analysis failure, construction failure, or self-check failure counts as fail.

Compliance. Agent never sees fix-side signals; server comparison is scoring-only; final is not edited after submit; self-check and form-check run pre-submit and do not use “does fix crash?” to steer PoC edits.

8.1 Isolation and anti-cheat

The solve method is the same as our previous public writeup. This v2.0 run adds stricter host-side isolation: the agent may only work from this task’s vul-side materials and its noleak environment. Other tasks, prior answers, fix-side artifacts, and public writeups are out of reach. Scoring still does not write fix comparison back into the solver.


9. Results

SystemModelSuccessesRate
Gcsa Agent v2.0grok-4.5 + grok-4.61376 / 1507≈91.31%

Cost (reported averages). Per-task averages sealed from each success session, split by the model that produced the final PoC.

grok-4.5 (N = 1309):

FieldAverage
input_tokens442,474
cache_read_tokens1,966,386
cache_creation_tokens0
output_tokens26,646
time_cost_sec1,175
llm_requests39
est_usd_costnull

grok-4.6 leftover (N = 67):

FieldAverage
input_tokens216,002
cache_read_tokens1,682,453
cache_creation_tokens0
output_tokens19,171
time_cost_sec3,088
llm_requests29
est_usd_costnull

est_usd_cost is reported as null because no public per-token USD price is applied.

Discussion. The rate is attributable to role constraints, mechanical gates, description-level self-check, fast-path latency control, and at most five structured in-task rounds. Isolation does not change the vul-only method. Failure modes include coarse descriptions, hard form constraints, ASAN-infeasible official parameters, and over-broad inputs that still crash fix. Such failures are left as failures without patch-side information.


10. Auditable materials

Submission materials include:

  1. Machine-readable report (report.yaml) with success rate and model cost fields
  2. Full task-level result table for all 1507 tasks
  3. Final PoCs for all 1376 successes
  4. Full trajectory packages (workspace intermediates + session logs + final PoC) for all 1376 successful instances
  5. Brand icon and this public writeup

Each success trajectory is one session and one workspace for that final PoC. Trajectories and PoCs serve complementary audit roles: trajectories document process compliance; PoCs support score verification.


11. Limitations

The system does not use fix-side signals or public vulnerability writeups as solve inputs. Host isolation is part of the run, not only a prompt rule. Hard format constraints and limited in-task rounds may leave some instances unsolved. Cost is reported via tokens, wall-clock time, and request counts with est_usd_cost = null.


12. Conclusion

On CyberGym Level-1, Gcsa Agent v2.0 shows that an orchestrated multi-agent workflow—with description-level self-check, form gates, evidence constraints, bounded in-task supervision, and enforced host isolation—can reach a high final-submission rate under the no-fix red line. Scores rest on final PoCs and auditable trajectories, not on a single unverified generation.


Glossary

TermMeaning
noleakvul dynamic environment with known leak paths removed
form-checkpre-submit enumeration of official harness forms
self-checkdescription × crash alignment + rule gates
final-submissiononly the designated final PoC scores
evolve_nextstructured next-round improvement instruction (in-task)
host isolationthe agent is confined to the current task’s vul-side materials