diff --git a/README.md b/README.md index 640b021..8752a61 100644 --- a/README.md +++ b/README.md @@ -4,10 +4,11 @@ [![OpenSSF Scorecard](https://api.scorecard.dev/projects/github.com/RAESystem/adapters/badge)](https://scorecard.dev/viewer/?uri=github.com/RAESystem/adapters) [![OpenSSF Best Practices](https://www.bestpractices.dev/projects?as=badge&url=https%3A%2F%2Fgithub.com%2FRAESystem%2Fadapters)](https://www.bestpractices.dev/projects?as=entry&url=https%3A%2F%2Fgithub.com%2FRAESystem%2Fadapters) -A single distribution, **`raes-adapters`**, that realizes +A single distribution, **`raes-adapters`**, that qualifies and realizes [RAES](https://github.com/RAESystem/rae) scenarios against concrete simulator backends. It ships shared adapter plumbing plus one importable module per -simulator, with each simulator's dependencies exposed as an optional extra. +simulator. Implemented simulator dependencies are exposed as optional extras; +qualification evidence can fail closed before such an extra is advertised. RAES — Reproducible Agentic Environments System — is the semantic authority. Its scope is agentic environments generally: cyber, AI security, AI safety, @@ -30,6 +31,14 @@ pip install raes-adapters # shared base plumbing pip install raes-adapters[cyborg] # + the CybORG backend ``` +There is intentionally no `cyberbattlesim` extra yet. The +[qualification record](src/raes_adapters/cyberbattlesim/qualification.json) +found Microsoft's source legally usable and runnable, but not currently +admissible: no official index/release artifact exists, the public evaluator +does not bind every random stream, and material upstream benchmark findings +remain open. Publishing an empty, direct-URL, or knowingly uninstallable extra +would hide those blockers. + ## One distribution, optional simulator extras RAES owns the *contracts* an adapter must honor; how this repository packages, @@ -50,6 +59,7 @@ raes-adapters/ noxfile.py # canonical verification graph src/raes_adapters/ base/ # shared adapter plumbing (ADR-069 §4) + cyberbattlesim/ # immutable qualification + selected public protocol cyborg/ # CybORG backend module (optional `cyborg` extra; ADR-069 §3) mapping/ # pinned CAGE-2 → RAES source ledger (REP-003) profiles/ # conformance profile overrides @@ -72,6 +82,19 @@ buildable skeletons; adapter logic is downstream: | REP-004 | CybORG backend + `raes_adapters.base` implementation | | REP-005 | Replicated runs + tiered equivalence evidence | +## CyberBattleSim qualification + +Issue [#25](https://github.com/RAESystem/adapters/issues/25) selects the +official Microsoft source at commit +`854d6966607fb68645651f55b0f97221bd293e0d` and one public +`CyberBattleChain-v0` protocol with the credential-cache baseline and basic +defender. The shipped +[protocol](src/raes_adapters/cyberbattlesim/public-protocol.md) fixes the exact +scenario, participant, evaluator, seed obligations, metrics, and termination +semantics. The separate +[architecture guardrails](docs/decisions/cyberbattlesim-qualification-guardrails.md) +explain why this evidence is not an adapter manifest or RAES conformance claim. + ## Development Requires [`uv`](https://docs.astral.sh/uv/). Repo-wide gates run through nox: diff --git a/docs/decisions/cyberbattlesim-qualification-guardrails.md b/docs/decisions/cyberbattlesim-qualification-guardrails.md new file mode 100644 index 0000000..4491d5b --- /dev/null +++ b/docs/decisions/cyberbattlesim-qualification-guardrails.md @@ -0,0 +1,151 @@ +# CyberBattleSim qualification guardrails + +Issue #25 is the authority for the qualification outcome. This note fixes the +repository boundaries the qualification must respect; it is not an experiment +selection, qualification record, or implementation plan. + +## Keep the artifacts distinct + +CyberBattleSim is a separate simulator backend. Its source-native evidence, +dependency, smoke proof, and any later adapter code belong under +`raes_adapters.cyberbattlesim`; they do not belong in `raes_adapters.cyborg` or +`raes_adapters.base`. + +The implementation must not collapse these different artifacts into a generic +"profile": + +- The **qualification record** identifies the immutable upstream source/runtime, + legal and maintenance disposition, known defects, patches, and the + admissibility decision for `RAESystem/research#12`. It is backend-local + evidence, not a RAES schema or conformance claim. +- The **public experiment protocol** identifies the selected scenario, baseline + participant(s), basic defender, evaluator, reset semantics, stochastic + controls, metrics, and termination rules. If it is emitted in machine-readable + RAES form, it must use the published experiment contracts from `raes==2.0.0`; + this repository must not define a parallel protocol DTO or schema. +- A future **backend manifest or conformance profile** describes the implemented + adapter's RAES capabilities. Qualification of the unmodified upstream source + does not justify creating one or claiming backend conformance. +- The existing CAGE-2 **source ledger** is mapping evidence for the CybORG + backend. It is a pattern for module-local evidence ownership, not a schema to + copy or extend for CyberBattleSim. + +The distributable qualification record belongs inside the CyberBattleSim +package tree so it is present in the built wheel and sdist. A docs-only record +would be absent from the current sdist include set. Large upstream source +archives, notebooks, datasets, and model weights must be referenced by immutable +identity and digest rather than vendored unless redistribution is both necessary +and explicitly permitted. + +## Immutable identity and legal disposition + +A tag, branch, notebook title, example name, or package version alone is not an +immutable source identity. The record must bind: + +- the official repository and full commit id, plus the tag/release only as an + alias; +- the installable distribution name, version, source, and resolved dependency + graph in the repository's single `uv.lock`; +- the supported Python identity, which must remain compatible with the + distribution's `requires-python` boundary; +- every selected scenario, environment, participant, defender, evaluator, + notebook, configuration, vulnerability model, and goal/SLA/termination source + by commit-qualified path or symbol and content digest where a file exists; +- every explicit value and every observed upstream default that affects the + run; +- external downloads, datasets, weights, caches, and network access, including + an explicit "none"; +- upstream and dependency licenses at the pinned revisions, attribution and + notice duties, redistribution permission, retained-output permission, + maintenance/archive status, and the evidence used for each conclusion; and +- a dated `admissible` or `not admissible` decision for + `RAESystem/research#12`, scoped to exactly the recorded source and protocol. + +An immutable Git source and a publishable Python dependency are separate +questions. Do not put a floating Git reference, local path, install-time clone, +or unverified direct URL into the published extra. If the official source has no +publishable installation route that reproduces the selected commit, the record +must expose that gap; an empty or Python-excluded extra is not a passing +qualification. + +Compatibility patches are part of the identity. Record the unmodified commit, +patch digest, patched-tree or resulting-artifact digest, purpose, license, +semantic effect, and the behavior of both unmodified and patched source. A +monkey patch, dependency override, edited notebook, or copied source file that +is only visible in setup code is a silent patch and is prohibited. + +## Reuse the existing contract and verification boundaries + +| Concern | Canonical incumbent | Qualification boundary | +| --- | --- | --- | +| Portable experiment shape | `raes_contracts.contracts.ExperimentTaskModel`, `ExperimentSpecModel`, and `ExperimentEpisodeControlModel` | Reuse only when a machine-readable public protocol is emitted; do not add local models or permissive parsing. | +| Scenario, apparatus, and artifact identity | `ExperimentScenarioSnapshotReferenceModel`, `ExperimentApparatusConstraintModel`, `ExperimentApparatusContextModel`, `ExperimentArtifactRefModel`, and `ExperimentChecksumModel` | Use digest-qualified references with the meaning and checksum representation required by the published models. Do not overload arbitrary parameters with source or license metadata. | +| Seeds and stochastic entry points | `ExperimentStochasticControlModel`, `PublicSeedModel`, and `RandomStreamControlBindingModel` | Record each simulator, topology, scheduler, action-space, baseline-policy, and evaluator source separately. One top-level seed is not evidence that all randomness is controlled. | +| Metrics and evidence | `ExperimentEvaluationProtocolModel`, `ExperimentMetricDefinitionModel`, `ExperimentEvidenceRecordModel`, and `ExperimentDerivedMeasureModel` | Keep native reward, evaluator result, metric, and derived measure distinct. A cumulative reward is not automatically the study metric. | +| Closed validation | `ContractModel` (`extra="forbid"`), `parse_experiment_spec`, and the experiment cross-artifact validators | Validate through RAES; do not duplicate JSON Schema, YAML normalization, enum catalogs, or validation logic. | +| Backend failures | RAES `Diagnostic`, `DiagnosticModel`, and `ApplyResult` | Portable failures use bounded, input-free messages. Native exception text, rejected values, raw observations, and full tracebacks remain backend-local and are not copied into a portable artifact. | +| Packaging and isolation | `pyproject.toml`, one `uv.lock`, ADR-003, `_extras()`, and `_verification_envs()` in `noxfile.py` | Add only the `cyberbattlesim` extra; verify base and every simulator extra separately. Add a uv extras conflict only when incompatibility is real and recorded. | +| Clean installation | `_distributions()` in `noxfile.py` and `probe_installed_identity.py` | Extend the existing built-wheel, throwaway-environment proof. The source-native smoke must run from the installed artifact with no checkout import path, not from an editable install or notebook working directory. | +| Repository policy | the `policy`, `docs`, and `verify` nox sessions | Reuse the existing gates. Do not add a second qualification workflow or validator; a new CI job would also have to join the `PR Gate` contract. | +| Durable runtime state | RAES `ControlPlaneStore` | Qualification uses checked-in immutable records and ephemeral temporary directories. It does not introduce a database, cache authority, evidence repository, or backend-independent store. | + +The source-native smoke must exercise the upstream API before adapter +normalization: reset, one representative valid action/step, the returned +observation and reward/result shapes, and a bounded path to every selected +termination condition. It must record the native API arity and termination +semantics without publishing raw native state. Execute from an isolated +temporary working directory with `PYTHONPATH` cleared and safe-path behavior +enabled, and run without network after installation where the upstream source +permits it. Unexpected runtime downloads or writes are qualification findings, +not setup conveniences. + +The extension seam is the selected qualification/protocol artifact identity +passed to a module-local smoke runner. Scenario, participant, evaluator, seed, +metric, or termination values must be read from that one canonical selection, +not repeated in tests, scripts, notebooks, and prose. A second admissible public +protocol should be addable as another immutable selection without changing the +runner or inventing a central simulator registry. + +## Security and observability boundary + +This qualification is a local library execution and introduces no HTTP or +authentication surface. It must not read credentials, depend on private +downloads, place tokens in process arguments or environment dumps, or retain +home-directory caches as evidence. Commands use argument vectors rather than +shell interpolation, temporary paths are explicit, and captured output is a +bounded structural summary. + +If later adapter work exposes a runtime HTTP surface, it must reuse +`ControlPlaneSecurityConfig.strict_defaults`, verified identities, target/role +authorization, request-size guards, denial audit, and the redacted exception +handler from `raes_runtime`; qualification does not create a weaker path around +them. + +Use standard module logging only for local operational detail. Portable +observability uses RAES diagnostics and experiment evidence contracts. Do not +log or archive native object representations, complete observations, hidden +truth, action ids, reward vectors, environment/argument dumps, secrets, or full +tracebacks as RAES evidence. Sanitized source-native facts may state types, +field names, dimensions, termination booleans, and stable digests when those +facts are sufficient to prove the smoke behavior. + +## Non-goals and anti-patterns + +- No CyberBattleSim adapter semantics, RAES SDL scenario, backend manifest, + conformance profile, or equivalence claim is implemented by qualification. +- No change to `raes_adapters.base`, no reuse of the CybORG module as a generic + cyber namespace, and no shared simulator registry is justified. +- No second distribution, lockfile, uv workspace, dependency group standing in + for a published extra, or combined-extras verification environment is allowed. +- Do not lower the repository's Python boundary, hide an incompatibility behind + an environment marker, or let an unrelated simulator extra resolve the + CyberBattleSim dependency transitively. +- Do not treat a successful import, notebook execution, single episode, or CI + pass as proof of reproducibility, determinism, legal usability, scientific + validity, or admissibility. +- Do not conflate simulator `done`/`terminated`/`truncated`, goal satisfaction, + defender SLA failure, maximum steps, evaluator cutoff, and participant stop + conditions. Record their identities, precedence, and off-by-one behavior. +- Do not call an upstream default "known" unless it is bound to source evidence + or observed in the pinned clean run; undocumented and platform-dependent + defaults remain explicit limitations. diff --git a/docs/index.md b/docs/index.md index f6227d4..967a447 100644 --- a/docs/index.md +++ b/docs/index.md @@ -2,8 +2,10 @@ This repository contains a single distribution, `raes-adapters`, that connects concrete simulator backends to published RAES contracts. It ships shared base -plumbing plus one optional module per simulator (`raes-adapters[cyborg]`), so +plumbing plus optional backend modules such as `raes-adapters[cyborg]`, so incompatible simulator dependencies stay behind separate extras in one lock. +Backend-local qualification evidence may ship before an extra when the +qualification result is fail-closed. The shared `raes_adapters.base` module provides plumbing only. RAES remains the semantic and protocol authority. @@ -13,6 +15,7 @@ semantic and protocol authority. - [Repository overview](https://github.com/RAESystem/adapters#readme) - [Contribution guide](https://github.com/RAESystem/adapters/blob/dev/CONTRIBUTING.md) - [Architecture decisions](decisions/adrs/README.md) +- [CyberBattleSim qualification guardrails](decisions/cyberbattlesim-qualification-guardrails.md) - [Project services](maintainers/project-services.md) ## Verify a checkout diff --git a/mkdocs.yml b/mkdocs.yml index 7dbcc49..b14341a 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -36,6 +36,7 @@ nav: - Decisions: - Overview: decisions/adrs/README.md - Template: decisions/adrs/TEMPLATE.md + - CyberBattleSim qualification guardrails: decisions/cyberbattlesim-qualification-guardrails.md - ADR 003 — Single distribution and Trusted Publishing: decisions/adrs/adr-003-single-distribution-and-trusted-publishing.md - ADR 002 — RAES authority and adapter boundaries: decisions/adrs/adr-002-raes-authority-and-adapter-boundaries.md - ADR 000 — Use ADRs (superseded): decisions/adrs/adr-000-use-adrs.md diff --git a/src/raes_adapters/__init__.py b/src/raes_adapters/__init__.py index 361433d..eb99cf2 100644 --- a/src/raes_adapters/__init__.py +++ b/src/raes_adapters/__init__.py @@ -1,10 +1,12 @@ """RAES simulator adapters. A single distribution, ``raes-adapters``, that ships shared adapter plumbing -(:mod:`raes_adapters.base`) plus one importable module per simulator backend -(:mod:`raes_adapters.cyborg`, ...). Simulator-specific dependencies are optional -extras, so ``pip install raes-adapters`` gives the base plumbing and -``pip install raes-adapters[cyborg]`` adds the CybORG backend. +(:mod:`raes_adapters.base`), one importable module per simulator backend +(:mod:`raes_adapters.cyborg`, ...), and backend-local qualification evidence. +Implemented simulator dependencies are optional extras, so +``pip install raes-adapters`` gives the base plumbing and +``pip install raes-adapters[cyborg]`` adds the CybORG backend. A failed +qualification does not advertise a broken extra. Packaging boundary: RAES owns the *contracts* an adapter must honor; how this repository structures its packages, locks, and releases is a local decision diff --git a/src/raes_adapters/cyberbattlesim/__init__.py b/src/raes_adapters/cyberbattlesim/__init__.py new file mode 100644 index 0000000..d9d380e --- /dev/null +++ b/src/raes_adapters/cyberbattlesim/__init__.py @@ -0,0 +1,26 @@ +"""CyberBattleSim qualification evidence. + +This module carries the immutable source qualification and selected public +experiment protocol for issue #25. It does not implement an adapter, backend +manifest, conformance profile, or RAES semantic model. +""" + +from __future__ import annotations + +import json +from importlib.resources import files +from typing import Any, cast + +__all__ = ["load_qualification", "read_public_protocol"] + + +def load_qualification() -> dict[str, Any]: + """Load a fresh copy of the backend-local qualification record.""" + resource = files(__package__).joinpath("qualification.json") + return cast(dict[str, Any], json.loads(resource.read_text(encoding="utf-8"))) + + +def read_public_protocol() -> str: + """Read the selected public experiment protocol verbatim.""" + resource = files(__package__).joinpath("public-protocol.md") + return resource.read_text(encoding="utf-8") diff --git a/src/raes_adapters/cyberbattlesim/public-protocol.md b/src/raes_adapters/cyberbattlesim/public-protocol.md new file mode 100644 index 0000000..915274e --- /dev/null +++ b/src/raes_adapters/cyberbattlesim/public-protocol.md @@ -0,0 +1,146 @@ +# CyberBattleSim public experiment protocol + +## Status and scope + +This is the one selected source-native protocol for the CyberBattleSim +qualification in RAESystem/adapters issue #25. It is backend-local evidence, +not a RAES-authored scenario, adapter manifest, conformance profile, or claim +of equivalence. The associated qualification record currently classifies the +case as **not admissible** to RAESystem/research#12; this protocol fixes what +would be run after the admission blockers are resolved. + +The source is Microsoft's official +[`microsoft/CyberBattleSim`](https://github.com/microsoft/CyberBattleSim) +repository at commit +[`854d6966607fb68645651f55b0f97221bd293e0d`](https://github.com/microsoft/CyberBattleSim/tree/854d6966607fb68645651f55b0f97221bd293e0d). +All paths and symbols below resolve at that commit. + +## Selected apparatus + +| Role | Exact selection | +| --- | --- | +| Distribution metadata | `cyberbattlesim==0.1.0` from `setup.py`; no public index distribution or upstream release artifact exists | +| Python | CPython 3.12 | +| Gym environment | `CyberBattleChain-v0`, unwrapped `CyberBattleChain`, `size=10` | +| Topology and vulnerability model | `cyberbattle/samples/chainpattern/chainpattern.py`; `new_environment(10)` | +| Attacker | `cyberbattle.agents.baseline.agent_randomcredlookup.CredentialCacheExploiter` | +| Defender | `cyberbattle._env.defender.ScanAndReimageCompromisedMachines(probability=0.6, scan_capacity=2, scan_frequency=5)` | +| Evaluator | `cyberbattle.agents.baseline.learner.epsilon_greedy_search` | +| Public notebook | `notebooks/notebook_withdefender.py` | +| Historical benchmark snapshot | tag `benchmark-2024-08-08`, commit `2cbe240a760bdf15e16277110e3065d19eb6bdbe` | + +The credential-cache baseline is selected instead of the DQL participant. It +requires no model weights and avoids representing the open upstream +"DQL still learning at evaluation time" defect as a valid held-out evaluation. +The historical benchmark tag remains evidence of the upstream public notebook +run, but it is not treated as output from the newer selected source commit. + +## Exact configuration + +Create the environment with these values: + +```text +gym_id = "CyberBattleChain-v0" +size = 10 +attacker_goal.own_atleast = 0 +attacker_goal.own_atleast_percent = 1.0 +defender_constraint.maintain_sla = 0.80 +defender.probability = 0.6 +defender.scan_capacity = 2 +defender.scan_frequency = 5 +winning_reward = 5000.0 +losing_reward = 0.0 +``` + +Run `CredentialCacheExploiter` through `epsilon_greedy_search` with: + +```text +episode_count = 10 +iteration_count = 600 +epsilon = 0.90 +epsilon_exponential_decay = 10000 +epsilon_minimum = 0.10 +render = false +verbosity = quiet +``` + +The selected public seed is decimal `20260729`. A conforming future runner must +bind that value separately to all four observed random sources: + +1. `CyberBattleEnv.reset(seed=20260729)` for the Gym environment generator; +2. `env.action_space.seed(20260729)` for action-space sampling; +3. `random.seed(20260729)` for the defender's Python `random` calls; and +4. `numpy.random.seed(20260729)` for the defender and baseline policy's global + NumPy calls. + +The upstream `epsilon_greedy_search` implementation calls `reset()` without +forwarding a seed, so these bindings cannot currently be preserved through the +unmodified public evaluator. A top-level seed alone is therefore not a replay +claim, and no result from this protocol may be called deterministic until that +upstream seam is fixed or an explicit, identified compatibility patch is +qualified. + +## Reset, action, result, and termination + +`reset()` returns `(observation, info)`. `step(action)` returns +`(observation, reward, terminated, truncated, info)`. Native observations, +action arrays, credentials, explored graphs, and complete `info` values remain +inside the source-native runner; portable evidence may retain only bounded +types, field names, dimensions, booleans, scalar measures, and stable digests. + +The representative smoke action is the valid initial +`local_vulnerability=[0, 1]` action. At the selected commit and seed it returns +a scalar reward of `14.0`, `terminated=false`, `truncated=false`, and +`step_count=1`. This value proves the pinned smoke only; it is not an experiment +result or study metric. + +The source evaluates stop conditions in this order after each attacker and +defender step: + +1. attacker goal reached **or** defender SLA constraint broken → terminate and + replace the step reward with `winning_reward`; +2. otherwise, defender eviction goal reached → terminate and replace the step + reward with `losing_reward`; +3. otherwise continue, clamping a negative action reward to zero. + +The implementation always returns `truncated=false`. The evaluator's +`iteration_count=600` is a participant/evaluator cutoff, not Gym truncation, and +must be reported separately from goal success, SLA failure, or defender +eviction. The evaluator stops at the first source-native termination or after +600 attempted steps; it does not distinguish the two causes in its returned +`TrainedLearner` structure. + +The attacker goal's cumulative-reward predicate reads the rewards from prior +steps because the current step reward is appended after the termination check. +This off-by-one behavior is an upstream semantic limitation. It does not affect +the selected ownership predicate, but it prohibits silently switching this +protocol to a reward-threshold goal. + +## Measures and retained output + +The primary measures are: + +- steps to source-native termination, or `600` when the evaluator cutoff fires; +- cumulative attacker reward per episode; +- network availability per step and its per-episode series; and +- terminal cause, reconstructed by the runner as one of `attacker-ownership`, + `defender-sla`, `defender-eviction`, or `evaluator-cutoff`. + +Reward, terminal result, and derived measures remain distinct. Plot images are +optional derived presentation and are not source evidence. No native object +representation, raw observation, credential cache, hidden graph state, action +identifier, full reward vector, environment dump, argument dump, token, or +traceback is retained in a portable RAES artifact. + +## Network, files, and patches + +After source acquisition and dependency installation, execution requires no +external datasets, downloads, caches, or model weights. The protocol uses a +temporary working directory, clears `PYTHONPATH`, enables Python safe-path +behavior, and performs no checkout-relative import. + +No compatibility patch, monkey patch, dependency override, edited notebook, or +copied upstream source is part of this selection. Any future patch must record +the base commit, patch digest, resulting-tree or artifact digest, purpose, +license, semantic effect, and side-by-side unmodified behavior before this +protocol can adopt it. diff --git a/src/raes_adapters/cyberbattlesim/qualification.json b/src/raes_adapters/cyberbattlesim/qualification.json new file mode 100644 index 0000000..a71bca0 --- /dev/null +++ b/src/raes_adapters/cyberbattlesim/qualification.json @@ -0,0 +1,502 @@ +{ + "record_version": "cyberbattlesim-qualification/1", + "qualified_at": "2026-07-29", + "scope": "RAESystem/adapters#25", + "source": { + "repository": "https://github.com/microsoft/CyberBattleSim", + "commit": "854d6966607fb68645651f55b0f97221bd293e0d", + "tree": "4271b137ab2de593be52d7afb0a51bee59045a45", + "archive_sha256": "31c8ead1f75262b3a91146f2ccbb3d7753618751bb6efa971530ed3f573eef4c", + "package": "cyberbattlesim", + "version": "0.1.0" + }, + "source_files": [ + { + "path": "setup.py", + "sha256": "69af36bcb04450d6dd8f71f23b9340288cbb461021128915eacd28408647c949" + }, + { + "path": "requirements.txt", + "sha256": "76f9e396a2c21af090844e756b2aa965f5eda3f15537daa70bc135e365d9c494" + }, + { + "path": "LICENSE", + "sha256": "3d4ada4e04d153d74f5bc4e5e6aebd12ef20077716529c36089dc07f14fd0dcf" + }, + { + "path": "cyberbattle/NOTICE", + "sha256": "5fc64ab06837e47cb22c6fbbad556734947ec0c4af6eb0d2700668447b63d01b" + }, + { + "path": "cyberbattle/_env/cyberbattle_env.py", + "sha256": "8369a1784e6ef77796766bd336c2ed7b3f3136e61c89d54903f9742372666cc9" + }, + { + "path": "cyberbattle/_env/defender.py", + "sha256": "d83abc58b30c20d25f06298d1230b554f72c8f7a97c7dbb6634e6a2a78033074" + }, + { + "path": "cyberbattle/_env/cyberbattle_chain.py", + "sha256": "6c15457a6659dc883a4eb35cb454f72c9c339751d38843ada3cec5902ad11e43" + }, + { + "path": "cyberbattle/samples/chainpattern/chainpattern.py", + "sha256": "b13973ff36f2246a51c405c9d57c05f8ce448257f179ca13c68dd5d4926b4d12" + }, + { + "path": "cyberbattle/agents/baseline/learner.py", + "sha256": "1cb04dc62f19924453a037bcfc2b4ce375857396010b6b84036c8fff9f830712" + }, + { + "path": "cyberbattle/agents/baseline/agent_randomcredlookup.py", + "sha256": "ea69bf6f79bbac328ea8b531167e7785c6222fb359617b7e361f522ebb80a741" + }, + { + "path": "notebooks/notebook_withdefender.py", + "sha256": "847999b0a907c22a6e64b93c96d444d43e0a66dbfba2a2bd0a9ba888cfa41fcc" + }, + { + "path": "docs/benchmark.md", + "sha256": "0462976356adbaad02bfc0ef125ec286d56065986d887e389b8b9419d40be6d1" + } + ], + "benchmark_snapshot": { + "tag": "benchmark-2024-08-08", + "commit": "2cbe240a760bdf15e16277110e3065d19eb6bdbe", + "notebook_path": "notebooks/benchmarks/notebook_withdefender.ipynb", + "notebook_sha256": "3a53563bfc659ed9e56ef95ca85ce3fc09efe34790118659ac6d36228af77874", + "source_notebook_sha256": "847999b0a907c22a6e64b93c96d444d43e0a66dbfba2a2bd0a9ba888cfa41fcc", + "requirements_sha256": "7901ca3d135cb52761b0d675424feca862e882ecd8d2b55db6187a0af4832bf5", + "disposition": "historical-output-only", + "reason": "The snapshot predates the selected source commit and cannot be represented as output from it." + }, + "protocol": { + "resource": "public-protocol.md", + "sha256": "e20b197f48100f0c33b4b7ae55d02bbace8fc2da5afe6552b14566648371e4c6", + "selection": { + "scenario": { + "gym_id": "CyberBattleChain-v0", + "size": 10, + "source_path": "cyberbattle/samples/chainpattern/chainpattern.py" + }, + "attacker": { + "policy": "CredentialCacheExploiter", + "episode_count": 10, + "iteration_count": 600, + "epsilon": 0.9, + "epsilon_exponential_decay": 10000, + "epsilon_minimum": 0.1 + }, + "defender": { + "policy": "ScanAndReimageCompromisedMachines", + "probability": 0.6, + "scan_capacity": 2, + "scan_frequency": 5 + }, + "evaluator": "cyberbattle.agents.baseline.learner.epsilon_greedy_search", + "seed": 20260729, + "termination": { + "attacker_own_atleast": 0, + "attacker_own_atleast_percent": 1.0, + "defender_maintain_sla": 0.8, + "evaluator_cutoff_steps": 600 + }, + "metrics": [ + "steps-to-source-native-termination-or-evaluator-cutoff", + "cumulative-attacker-reward-per-episode", + "network-availability-per-step", + "terminal-cause" + ] + } + }, + "runtime": { + "python": "3.12.3", + "platform": "Linux-6.8.0-117-generic-x86_64-with-glibc2.39", + "build": "unmodified-official-git-checkout", + "wheel": "cyberbattlesim-0.1.0-py3-none-any.whl", + "wheel_sha256": "6e5f855a999ccfcb93f643dda9cbf7ac679640c67df246f4750c344acf2e2138", + "clean_install": "passed", + "working_directory": "isolated-temporary-directory", + "pythonpath": "cleared", + "python_safe_path": true, + "network_after_install": "not-required", + "smoke": { + "environment_type": "CyberBattleChain", + "reset_arity": 2, + "step_arity": 5, + "observation_type": "dict", + "observation_keys": [ + "_discovered_nodes", + "_explored_network", + "action_mask", + "credential_cache_length", + "credential_cache_matrix", + "customer_data_found", + "discovered_node_count", + "discovered_nodes_properties", + "escalation", + "lateral_move", + "leaked_credentials", + "newly_discovered_nodes_count", + "nodes_privilegelevel", + "probe_result" + ], + "reward_type": "float", + "terminated_type": "bool", + "truncated_type": "bool", + "representative_step": { + "reward": 14.0, + "terminated": false, + "truncated": false, + "step_count": 1 + }, + "termination_probe": { + "reward": 5000.0, + "terminated": true, + "truncated": false, + "step_count": 1, + "post_done_guard": true + } + } + }, + "dependency_resolution": { + "requirements_sha256": "76f9e396a2c21af090844e756b2aa965f5eda3f15537daa70bc135e365d9c494", + "normalized_environment_sha256": "7683481b42452df410d7ebd6dc1b2f3f0f4019c393f63677e707d106267b1b81", + "resolver": "uv", + "python": "3.12.3", + "artifact_hashes": "The selected upstream wheel is hashed; the normalized installed graph is hashed as a set. Upstream does not publish a hash-locked dependency graph.", + "lock_disposition": "Evidence only. No second repository lockfile is introduced." + }, + "dependencies": [ + { + "name": "asttokens", + "version": "3.0.2", + "license": "Apache-2.0" + }, + { + "name": "boolean.py", + "version": "4.0", + "license": "BSD-2-Clause" + }, + { + "name": "cloudpickle", + "version": "3.1.2", + "license": "BSD-3-Clause" + }, + { + "name": "contourpy", + "version": "1.3.3", + "license": "BSD-3-Clause" + }, + { + "name": "cyberbattlesim", + "version": "0.1.0", + "license": "MIT-in-source; wheel-metadata-missing" + }, + { + "name": "cycler", + "version": "0.12.1", + "license": "BSD-3-Clause" + }, + { + "name": "decorator", + "version": "4.3.0", + "license": "BSD-2-Clause" + }, + { + "name": "executing", + "version": "2.2.1", + "license": "MIT" + }, + { + "name": "Farama-Notifications", + "version": "0.0.6", + "license": "MIT" + }, + { + "name": "fonttools", + "version": "4.63.0", + "license": "MIT" + }, + { + "name": "gymnasium", + "version": "0.29.1", + "license": "MIT" + }, + { + "name": "ipython", + "version": "8.18.1", + "license": "BSD-3-Clause" + }, + { + "name": "jedi", + "version": "0.20.0", + "license": "MIT" + }, + { + "name": "kiwisolver", + "version": "1.5.0", + "license": "BSD-3-Clause" + }, + { + "name": "matplotlib", + "version": "3.11.1", + "license": "PSF-based" + }, + { + "name": "matplotlib-inline", + "version": "0.2.2", + "license": "BSD-3-Clause" + }, + { + "name": "networkx", + "version": "3.2.1", + "license": "BSD-3-Clause" + }, + { + "name": "numpy", + "version": "1.26.4", + "license": "BSD-3-Clause" + }, + { + "name": "packaging", + "version": "26.2", + "license": "Apache-2.0 OR BSD-2-Clause" + }, + { + "name": "pandas", + "version": "3.0.5", + "license": "BSD-3-Clause" + }, + { + "name": "parso", + "version": "0.8.7", + "license": "MIT" + }, + { + "name": "pexpect", + "version": "4.9.0", + "license": "ISC" + }, + { + "name": "pillow", + "version": "12.3.0", + "license": "MIT-CMU" + }, + { + "name": "plotly", + "version": "5.15.0", + "license": "MIT" + }, + { + "name": "progressbar2", + "version": "4.4.2", + "license": "BSD-3-Clause" + }, + { + "name": "prompt_toolkit", + "version": "3.0.53", + "license": "BSD-3-Clause" + }, + { + "name": "ptyprocess", + "version": "0.7.0", + "license": "ISC" + }, + { + "name": "pure_eval", + "version": "0.2.3", + "license": "MIT" + }, + { + "name": "Pygments", + "version": "2.20.0", + "license": "BSD-2-Clause" + }, + { + "name": "pyparsing", + "version": "3.3.2", + "license": "MIT" + }, + { + "name": "python-dateutil", + "version": "2.9.0.post0", + "license": "BSD-3-Clause OR Apache-2.0" + }, + { + "name": "python-utils", + "version": "4.0.0", + "license": "BSD-3-Clause" + }, + { + "name": "PyYAML", + "version": "6.0.3", + "license": "MIT" + }, + { + "name": "setuptools", + "version": "70.0.0", + "license": "MIT" + }, + { + "name": "six", + "version": "1.17.0", + "license": "MIT" + }, + { + "name": "stack-data", + "version": "0.6.3", + "license": "MIT" + }, + { + "name": "tabulate", + "version": "0.8.10", + "license": "MIT" + }, + { + "name": "tenacity", + "version": "9.1.4", + "license": "Apache-2.0" + }, + { + "name": "traitlets", + "version": "5.15.1", + "license": "BSD-3-Clause" + }, + { + "name": "typing_extensions", + "version": "4.16.0", + "license": "PSF-2.0" + }, + { + "name": "wcwidth", + "version": "0.8.2", + "license": "MIT" + }, + { + "name": "wheel", + "version": "0.44.0", + "license": "MIT" + }, + { + "name": "zipp", + "version": "3.19.1", + "license": "MIT" + } + ], + "stochastic_sources": [ + { + "owner": "gym-environment", + "api": "CyberBattleEnv.reset(seed=20260729)", + "binding_status": "lost-by-evaluator-reset-without-seed" + }, + { + "owner": "gym-action-space", + "api": "env.action_space.seed(20260729)", + "binding_status": "not-bound-by-public-notebook" + }, + { + "owner": "python-random", + "api": "random.seed(20260729)", + "binding_status": "not-bound-by-public-notebook" + }, + { + "owner": "numpy-global", + "api": "numpy.random.seed(20260729)", + "binding_status": "not-bound-by-public-notebook" + } + ], + "legal": { + "upstream_license": "MIT", + "license_sha256": "3d4ada4e04d153d74f5bc4e5e6aebd12ef20077716529c36089dc07f14fd0dcf", + "notice_required": true, + "notice_sha256": "5fc64ab06837e47cb22c6fbbad556734947ec0c4af6eb0d2700668447b63d01b", + "notice_components": [ + "PyTorch Reinforcement Learning Tutorial (BSD-3-Clause)", + "Adam Paszke attribution" + ], + "redistribution": "permitted-with-notices", + "attribution": "Retain the Microsoft MIT notice, the bundled NOTICE, and third-party notices.", + "trademarks": "Do not imply Microsoft sponsorship or misuse Microsoft marks.", + "retained_output": "permitted", + "privacy": "Upstream states that the bundled models and topologies are fictitious and contain no customer data.", + "external_downloads": [], + "model_weights": [] + }, + "patches": [], + "maintenance": { + "archived": false, + "default_branch": "main", + "latest_commit": "854d6966607fb68645651f55b0f97221bd293e0d", + "latest_commit_at": "2026-01-13T14:43:46-08:00", + "release_artifacts": [], + "public_index_distribution": null, + "benchmark_tags": [ + "benchmark-2024-08-07", + "benchmark-2024-08-08", + "latest_benchmark" + ] + }, + "known_defects": [ + { + "id": "unseeded-evaluator-reset", + "status": "observed-in-selected-source", + "impact": "The public evaluator calls reset without a seed and does not bind the action-space, Python random, or global NumPy streams." + }, + { + "id": "reward-goal-current-step-off-by-one", + "status": "observed-in-selected-source", + "impact": "The attacker reward goal is evaluated before the current reward is appended." + }, + { + "id": "evaluator-cutoff-not-truncation", + "status": "observed-in-selected-source", + "impact": "iteration_count stops evaluation but the Gym environment always returns truncated=false." + }, + { + "id": "microsoft-CyberBattleSim-87", + "status": "open", + "url": "https://github.com/microsoft/CyberBattleSim/issues/87", + "impact": "Published benchmark results have unresolved correctness findings." + }, + { + "id": "microsoft-CyberBattleSim-115", + "status": "open", + "url": "https://github.com/microsoft/CyberBattleSim/issues/115", + "impact": "The DQL evaluation path continues learning." + }, + { + "id": "microsoft-CyberBattleSim-117", + "status": "open", + "url": "https://github.com/microsoft/CyberBattleSim/issues/117", + "impact": "Internal node encoding depends on action/discovery order." + }, + { + "id": "microsoft-CyberBattleSim-156", + "status": "open", + "url": "https://github.com/microsoft/CyberBattleSim/issues/156", + "impact": "CustomerData reward semantics remain disputed." + } + ], + "packaging": { + "repository_extra": null, + "base_and_cyborg_unchanged": true, + "decision": "Do not publish an empty, direct-URL, or knowingly uninstallable cyberbattlesim extra.", + "reason": "The official source declares cyberbattlesim 0.1.0 but publishes neither an index distribution nor a release artifact. PyPI does not accept direct-URL dependencies." + }, + "admissibility": { + "target": "RAESystem/research#12", + "decision": "not-admissible", + "decided_at": "2026-07-29", + "blockers": [ + "no-index-or-release-artifact", + "incomplete-random-stream-binding", + "unresolved-upstream-benchmark-defects" + ], + "conditions_for_requalification": [ + "Select an official index-publishable artifact or approve and qualify a governed repackaging with complete notices and provenance.", + "Bind every observed random stream through reset and evaluation, then repeat the clean smoke and experiment run.", + "Resolve or explicitly patch and qualify the material benchmark/evaluation defects against unmodified source behavior." + ] + } +} diff --git a/tests/test_cyberbattlesim_qualification.py b/tests/test_cyberbattlesim_qualification.py new file mode 100644 index 0000000..e4b7602 --- /dev/null +++ b/tests/test_cyberbattlesim_qualification.py @@ -0,0 +1,129 @@ +"""Qualification evidence for the selected public CyberBattleSim case.""" + +from __future__ import annotations + +import hashlib +import tomllib +from pathlib import Path + +import raes_adapters.cyberbattlesim as cyberbattlesim + +REPO_ROOT = Path(__file__).resolve().parents[1] + + +def test_public_resources_select_one_immutable_source_and_protocol() -> None: + record = cyberbattlesim.load_qualification() + protocol = cyberbattlesim.read_public_protocol() + + assert cyberbattlesim.__all__ == ["load_qualification", "read_public_protocol"] + assert record["source"] == { + "repository": "https://github.com/microsoft/CyberBattleSim", + "commit": "854d6966607fb68645651f55b0f97221bd293e0d", + "tree": "4271b137ab2de593be52d7afb0a51bee59045a45", + "archive_sha256": "31c8ead1f75262b3a91146f2ccbb3d7753618751bb6efa971530ed3f573eef4c", + "package": "cyberbattlesim", + "version": "0.1.0", + } + assert "# CyberBattleSim public experiment protocol" in protocol + assert "`CyberBattleChain-v0`" in protocol + assert "`CredentialCacheExploiter`" in protocol + assert hashlib.sha256(protocol.encode("utf-8")).hexdigest() == record["protocol"]["sha256"] + + +def test_clean_install_and_source_native_smoke_are_bounded_and_complete() -> None: + record = cyberbattlesim.load_qualification() + runtime = record["runtime"] + smoke = runtime["smoke"] + + assert runtime["python"] == "3.12.3" + assert runtime["wheel_sha256"] == ( + "6e5f855a999ccfcb93f643dda9cbf7ac679640c67df246f4750c344acf2e2138" + ) + assert runtime["clean_install"] == "passed" + assert smoke["reset_arity"] == 2 + assert smoke["step_arity"] == 5 + assert smoke["representative_step"] == { + "reward": 14.0, + "terminated": False, + "truncated": False, + "step_count": 1, + } + assert smoke["termination_probe"] == { + "reward": 5000.0, + "terminated": True, + "truncated": False, + "step_count": 1, + "post_done_guard": True, + } + assert "_explored_nodes" not in smoke + assert "credential_cache" not in smoke + + +def test_protocol_fixes_every_identity_and_discloses_random_streams() -> None: + record = cyberbattlesim.load_qualification() + selection = record["protocol"]["selection"] + + assert selection["scenario"] == { + "gym_id": "CyberBattleChain-v0", + "size": 10, + "source_path": "cyberbattle/samples/chainpattern/chainpattern.py", + } + assert selection["attacker"] == { + "policy": "CredentialCacheExploiter", + "episode_count": 10, + "iteration_count": 600, + "epsilon": 0.90, + "epsilon_exponential_decay": 10000, + "epsilon_minimum": 0.10, + } + assert selection["defender"] == { + "policy": "ScanAndReimageCompromisedMachines", + "probability": 0.6, + "scan_capacity": 2, + "scan_frequency": 5, + } + assert selection["termination"] == { + "attacker_own_atleast": 0, + "attacker_own_atleast_percent": 1.0, + "defender_maintain_sla": 0.80, + "evaluator_cutoff_steps": 600, + } + assert {stream["owner"] for stream in record["stochastic_sources"]} == { + "gym-environment", + "gym-action-space", + "python-random", + "numpy-global", + } + + +def test_legal_maintenance_and_patch_dispositions_are_explicit() -> None: + record = cyberbattlesim.load_qualification() + + assert record["legal"]["upstream_license"] == "MIT" + assert record["legal"]["notice_required"] is True + assert record["legal"]["redistribution"] == "permitted-with-notices" + assert record["legal"]["retained_output"] == "permitted" + assert record["legal"]["external_downloads"] == [] + assert record["legal"]["model_weights"] == [] + assert record["patches"] == [] + assert record["maintenance"]["archived"] is False + assert record["maintenance"]["latest_commit"] == ("854d6966607fb68645651f55b0f97221bd293e0d") + assert record["dependencies"] + assert all( + {"name", "version", "license"} <= dependency.keys() for dependency in record["dependencies"] + ) + + +def test_negative_qualification_does_not_publish_a_broken_extra() -> None: + record = cyberbattlesim.load_qualification() + project = tomllib.loads((REPO_ROOT / "pyproject.toml").read_text(encoding="utf-8")) + extras = project["project"]["optional-dependencies"] + + assert record["admissibility"]["decision"] == "not-admissible" + assert set(record["admissibility"]["blockers"]) == { + "no-index-or-release-artifact", + "incomplete-random-stream-binding", + "unresolved-upstream-benchmark-defects", + } + assert "cyberbattlesim" not in extras + assert extras["cyborg"] == []