Entry-point doc for how GYMSIEGE's Daytona adapter fits together, starting from
the client object every entrypoint opens first. For issue-specific deep dives
(the ARVO sanitizer oracle, the 60-minute safety TTL, the nested-container
egress bug) see daytona-notes.md; this doc is the
narrower "what talks to what" map.
AsyncDaytona is the async client for the Daytona SDK used throughout the
GYMSIEGE codebase as the entry point for all sandbox lifecycle operations โ
anything that creates, finds, or enumerates sandboxes/snapshots rather than
acting on one you already hold. Specifically it's used to:
- Create sandboxes โ
daytona.create()for a warm-pool-eligible "cold" sandbox, ordaytona.create(CreateSandboxFromSnapshotParams(...))to spin one up from the pre-bakedgymsiege-toolchainsnapshot. The snapshot form is called directly insnapshot_build.py(baking the toolchain and taking a restore-latency sample) and insandbox_runner.py:_create_sandbox(every real CyberGym trial); the cold form is theprovisioning == "cold"arm in that same function, plus the baseline sample inorchestrator.py'sprovision-benchcommand.orchestrator.py runreachesdaytona.create()too, but indirectly โ throughsandbox_runner.run_trialfor ordinary trials, and directly for the--provisioning forkwarm-parent sandbox it creates once up front before fanning out.fork()calls against it.sweepnever callscreate()itself either; it drives the samerun_trialpath at increasing concurrency levels. - List sandboxes โ
daytona.list()(an async generator, not something youawaitinto a list โ see the note inorchestrator.py) in thereapcommand, to find and delete straysiege-*sandboxes left behind by a crashed run. - Fetch a sandbox by id โ
daytona.get(...)indashboard.py'spublish()function, to reuse an existing dashboard sandbox instead of recreating one on every redeploy. - Manage the snapshot catalog โ
daytona.snapshot.list()(returns aPaginatedSnapshotsobject with an.itemslist, not a bare list) shows up in two places:demo.sh's inline Python snippet, to check whethergymsiege-toolchainalready exists before rebaking it, andexploitgym_snapshot_build.py, doing the same existence check for thegymsiege-exploitgymsnapshot before that bake runs.
It's always opened as an async context manager
(async with AsyncDaytona() as daytona:), which reads
DAYTONA_API_KEY/DAYTONA_API_URL/DAYTONA_TARGET from the environment
(loaded from .env/.env.local via common.py's load_dotenv calls, so
those vars don't need to be exported by hand) and handles client teardown on
exit โ configure_secrets.py, dashboard.py, orchestrator.py,
snapshot_build.py, and exploitgym_snapshot_build.py all follow this same
async with pattern rather than manually opening/closing the client.
Every individual sandbox object AsyncDaytona returns (AsyncSandbox) then
exposes the per-sandbox operations โ process.exec, computer_use.*,
get_metrics/get_metrics_latest, update_network_settings, set_ttl,
delete, fork, create_snapshot, download_url, update_secrets โ that
the rest of the pipeline (sandbox_runner.py, solver_agent.py) uses to
actually run each CyberGym-E2E or ExploitGym trial once the sandbox exists.
AsyncDaytona hands one out; everything after that is a method call on that
one object until sandbox.delete() ends its life.
One AsyncSandbox method combination worth calling out on its own:
update_network_settings(network_block_all=True/False), used in
BuildAgent._reconfirm_isolated (solver_agent.py).
CyberGym's own scripts/run_agent.py needs network access for the whole
find-vuln โ PoC โ patch loop, since the solver's own LLM calls go out over
the network throughout โ so the sandbox can't be network-cut for that entire
phase. Instead, once run_agent.py has finished and frozen its artifacts
(poc.bin / fix.patch, located from the newest summary.json under
agent_output/<task>/), BuildAgent.run calls _reconfirm_isolated, which:
- Calls
sandbox.update_network_settings(network_block_all=True)โ the network is now cut at the Daytona sandbox level. - Runs a small generated Python script inside the sandbox
(
_isolated_oracle_script) that reuses CyberGym's own container helpers (start_container,setup_workspace,copy_to_container,exec_run,cleanup_containerfrom itsscripts/utils.py) to spin up two fresh, nested per-arm containers โ one for the vulnerable build, one withfix.patchapplied โ copies the frozenpoc.bininto each, and runsrun_poc.shdirectly to capture its raw exit code (run_poc_exit_code), alongside CyberGym's ownvalidate.pystage result, for each arm. Because the outer sandbox's network is already blocked at this point, a missing base image fails closed here instead of silently pulling one mid-detonation. - Parses the two raw exit codes out of a
GYMSIEGE_ORACLE_JSON:marker line in the script's stdout asvul_exit_code(unpatched build) andfix_exit_code(patched build) โ these are whatresults.jsonreports, distinct fromrun_agent.py's own network-attached internal stage1/stage2 validation. - Re-opens the network (
network_block_all=False) in afinallyblock regardless of outcome, so the subsequent artifactdownload_urlcalls, recording upload, and cleanup exec still work.
BuildResult.network_isolated_detonation is only True when both exit codes
came back non-None; a failure anywhere in that sequence is captured as
detonation_error on the result instead of silently dropping the trial.
The CyberGym dataset payload is fetched during the bake
(snapshot_build.py's BOOTSTRAP_SH), authenticated by the
gymsiege-huggingface Daytona Secret. That Secret is scoped to nine exact
FQDNs, defined once in configure_secrets.py as HUGGINGFACE_SECRET_HOSTS
and mirrored in HUGGINGFACE_HOSTS.md:
| # | Host | Role |
|---|---|---|
| 1 | huggingface.co |
Hub API โ auth, gated-repo resolution, metadata |
| 2 | cas-server.xethub.hf.co |
Xet content-addressed store, US |
| 3 | cas-server.xethub-eu.hf.co |
Xet content-addressed store, EU |
| 4 | transfer.xethub.hf.co |
Xet transfer endpoint, US |
| 5 | transfer.xethub-eu.hf.co |
Xet transfer endpoint, EU |
| 6 | us.aws.cdn.hf.co |
CDN payload delivery, AWS US โ the observed redirect target |
| 7 | us.gcp.cdn.hf.co |
CDN payload delivery, GCP US |
| 8 | cdn-lfs-us-1.hf.co |
Legacy LFS CDN, US |
| 9 | cdn-lfs-eu-1.hf.co |
Legacy LFS CDN, EU |
Two things this list is not. It is not a network allowlist โ a Daytona
Secret's hosts field is a value-substitution trust boundary governing where
Daytona may send that secret. And it is not permanent: it tracks Hugging Face's
current download-behind-a-firewall guidance, so re-review it whenever
huggingface_hub is upgraded. tests/test_core.py asserts the list is
duplicate-free and still contains the key members.
This list is not what blocks the bake. An A/B probe
(hf_header_probe.py) transferred payload bytes from us.aws.cdn.hf.co
inside a sandbox (ranged GET โ 206), retiring the egress theory. The
confirmed cause is that responses reaching a sandbox arrive with
Content-Length removed and Transfer-Encoding added โ a proxy re-framing
them as chunked. huggingface_hub takes a file's size from X-Linked-Size,
or from Content-Length only when the response is not a redirect
(file_download.py:1645-1648), so the bake's 20 plain-git crash.log files
(direct 200, no X-Linked-Size) abort while its 40 LFS/Xet-backed
src.tgz/poc.bin files redirect and survive. Details in
reference/DAYTONA_HUGGINGFACE_EGRESS_ISSUE.md.
Note also that this dataset is Xet-backed, so huggingface_hub's real payload
path is Xet rather than the CDN redirect (file_download.py:1777); Xet is
Hugging Face's current backend and Git LFS the legacy one, not the reverse.
Telemetry is captured at two points in a trial's life, and the second one exists specifically so a trial that dies early still leaves evidence behind.
- Normal path โ a trial that reaches the end of its build/PoC/patch
phase (including the isolated re-detonation above) calls
sandbox.get_metrics_latest()andsandbox.get_metrics(start=None, end=None)atsandbox_runner.py:210-216, right afterstatusis classified. The results land on theTrialResultasmetrics_latest(one point-in-time sample) andmetrics_series(the time-series across the sandbox's lifetime โcpu_used_pct,mem_used,mem_total,disk_used, etc., straight from the realSandboxMetricsthe SDK returns). Neither call is periodic polling:get_metrics()with no bounds just pulls back whatever history Daytona's backend already accumulated for that sandbox. - Cancellation/timeout path โ
run_trial'sfinallyblock calls_capture_telemetry_bounded(sandbox_runner.py:260, defined at line 299) before the sandbox is deleted. It re-runs the same two SDK calls under a 30-secondasyncio.wait_forand swallows every exception. So a trial cancelled mid-build by the sweep's outer deadline still gets telemetry, provided the control plane is responsive enough to answer. This capture is best-effort by design โ an unresponsive control plane logs a warning and leaves the fieldsNonerather than blocking cleanup. - Retaining the partial result โ the sweep's own wrapper,
orchestrator.py:_run_sweep_trial_with_timeout, pre-allocates theTrialResultand passes it intorun_trialas theresult=argument, so the object thefinallyblock mutates is the same one the wrapper returns on timeout. Previously the timeout arm returnedNoneand threw the telemetry away even when it had been fetched. - Classification โ once a probe batch at a given concurrency level
finishes,
orchestrator.py:_trial_hit_oom_threshold(line 598) walks every sample in that trial'smetrics_seriesplus its finalmetrics_latest, flagging the trial as OOM-adjacent if any single sample showsmem_used >= 0.95 * mem_total.n_oom(line 550) counts those, andoom_rate = n_oom / ngoes intoresults/concurrency_sweep.json.
So "OOM" here really means "this sandbox's real memory telemetry crossed 95% utilization at some point during the trial," not a captured OOM-killer event, exit code, or kernel log line โ Daytona's SDK doesn't expose one. It's a proxy: high enough that memory pressure plausibly contributed to whatever else went wrong at that concurrency level (a failed build, a killed process, a timeout), but it's correlational, not a confirmed cause. The 95% threshold is hardcoded (not configurable via CLI/env) โ worth knowing if you want to tune sensitivity.
Timeout and OOM now overlap on purpose. Because a timed-out trial can
carry telemetry, it can be counted in n_timeout and n_oom
simultaneously. That is intended: the two describe different things (how the
trial ended, versus what its memory was doing), and forcing them to be
disjoint is what previously hid memory pressure behind timeouts. The overlap
is reported explicitly as n_timeout_oom / timeout_oom_rate
(orchestrator.py:551-566), so timeout_rate + oom_rate should not be read
as a sum of distinct failures. A timeout whose telemetry fetch also failed
stays unclassified โ it is not assumed to be non-OOM.
Caveat โ ExploitGym has not been fixed. The above applies to CyberGym's
run_trial only. exploitgym_adapter.py still captures telemetry once, very
late (lines 524-527), after evaluation, the network block, and result.json
scoring โ so its error, cancellation, and outer-deadline paths all discard
telemetry exactly the way sweep used to. Nothing surfaces this yet, because
exploitgym-run is a flat semaphore fan-out with no concurrency ladder and
_write_exploitgym_results computes no OOM statistic at all. It becomes a
real data-loss bug the moment anyone adds either. Tracked as Priority 8 in
TODO.md.
Both paths still leave a gap for the dashboard: trials whose telemetry fetch
failed outright render nothing in the per-sandbox charts, which only draw
trials with a non-empty metrics_series.