A read-only, plain-text, governed security-operations "brain" built on top of my own multi-agent runtime, phantom-mesh. The thesis in one line: don't build the engine — build the brain. Wrap mature, battle-tested scanners (Trivy / Nmap / Nuclei / Sigma) and put the value in the agent layer that orchestrates, correlates, and explains their output. The LLM triages; it never exploits.
📖 Full bilingual deep-dive (positioning · roadmap · OSS landscape · engineering decisions): docs/phantom-secops.md | ⚖️ Legal / ethical boundaries: ETHICS.md
An attack pipeline (recon → vuln-scan → exploit-suggest → report) and a defense pipeline (log-anomaly → triage → correlate → incident-report) run in parallel against an isolated, intentionally-vulnerable lab, on two clocks that both start at t=0. The headline metric is the one a SOC actually measures — Mean Time To Detect (MTTD):
→ MTTD = 15s (defender win)
defender triaged the activity at t+15s; attacker reached impact at t+50s
→ detected 35s before impact
The metric is honest in both directions: if detection lands after impact, the report says so
(attacker win). In mock mode the per-step timing is explicitly labelled simulated; live mode uses
real wall-clock. Every run also writes machine-readable summary.json + interleaved timeline +
ATT&CK phases + P1/P2/P3 queue.
Read-only inspection of this machine — host posture, dependency CVEs, host intrusion detection — merged by a deterministic, no-LLM fusion step and then synthesized by an LLM agent into one prioritised, plain-language action list. Data never leaves the machine. A real run on my own box surfaced 864 fixable CVEs in a sister project plus an AV real-time-protection gap, and the agent returned exact upgrade versions in fix-first order.
.\checkup.ps1 # one shot: tests + every engine + deterministic fusion + AI report
.\checkup.ps1 -Path C:\path\to\your-app # scan a specific project for CVEsThis repo is published as an installable Python package for the hermetic read-only public demo surface, with source-checkout commands retained for the lab and verification workflows. The official public entrypoints are:
pip install -e .for local package installation from a checkout.phantom-secops --helpfor the installable top-level CLI.phantom-secops reasoning-scenario --out reports/reasoning-demofor the installable, hermetic, read-only public reasoning smoke.phantom-secops compare <run-a> <run-b>to diff two kill-chain reports or run directories (read-only): align the two timelines, cross-reference attacker findings against defender IoCs, and report the numeric deltas.python scenarios/run_kill_chain.py --target juice-shop --mockfor the CI-safe red/blue demo with no Docker and no API key.python -m pytest -qfor the deterministic test suite.python scripts/run_goal.py --out reports/verificationfor the verification pack.python -m phantom_secops.defensive_loop --out reports/defensive-demofor the hermetic defensive finding/timeline artifact loop.python -m phantom_secops.evidence_playbook --out reports/evidence-playbookfor the hermetic evidence pack and tabletop playbook simulation..\checkup.ps1on Windows for the local endpoint self-check.make ...targets on Unix-like shells as shortcuts over the Python scripts.
Live lab targets and mesh/LLM paths are opt-in. The default public demo path is mock/read-only and does not scan external systems.
# Demo 1 — red/blue kill-chain + MTTD (~1s, no Docker, no API key)
pip install -e .
phantom-secops --help
phantom-secops reasoning-scenario --out reports/reasoning-demo
make demo-mock # or: python scenarios/run_kill_chain.py --target juice-shop --mock
# Demo 1 (live) — real nmap + nuclei against the Docker lab
make lab-up && make demo && make lab-down
# Goal verification pack (mock kill-chain + checkup + governance smoke + model diff)
make verify-goal
# Diff two kill-chain runs (read-only): align timelines, cross-ref findings <-> IoCs, report deltas
phantom-secops compare reports/verification/<ts>/killchain reports/verification/<other-ts>/killchain
# P2 hermetic defensive workbench loop: no active scanning, no Docker, no PoC
python -m phantom_secops.defensive_loop --out reports/defensive-demo
# P2 evidence pack + playbook simulation: metadata-only, no actions executed
python -m phantom_secops.evidence_playbook --out reports/evidence-playbook
# Strict audit (governance required + cross-model parity required)
make verify-goal-strict
# Optional cross-model / mesh smoke: set GOAL_MODEL_RUNNER_<MODEL> to a command
# template (see the Makefile) to diff kill-chain output across codex/claude/hermes,
# then `make verify-goal-strict` (or `make verify-goal-mesh` for the mesh path).
# Verify — full deterministic test suite (no real scanning in CI)
python -m pytest -qArtifacts land in reports/verification/<ts>/:
killchain/baseline run:incident-report.md,pentest-report.md,summary.jsoncheckup/checkup.txt: checkup text artifactgovernance/governance decisions:governance.jsonlandgovernance.loggoal-manifest.jsonwith stable fields:run(mesh+provider)killchain_signature/checkup_signaturemodelsoutcomesmodel_compare(killchainandcheckup)governanceblock (decision_count,decision_values, audit file)
audit_summarywithkillchain.mttd / outcome / detect_marginandgovernance_log
P2 defensive-loop artifacts land in the directory passed to
python -m phantom_secops.defensive_loop --out <dir>:
findings.jsonlwith schema version 1 defensive findings.timeline.jsonwith checkup -> verify -> analyze events.analysis.jsonwith ranked defensive actions.verification.jsonwith no-active-scan and no-runnable-PoC checks.manifest.jsonrecordingactive_scanning=false,external_network=false,exploit_poc=false, andwrites_to_host=false.
P2 evidence/playbook artifacts land in the directory passed to
python -m phantom_secops.evidence_playbook --out <dir>:
evidence-pack.jsonwith metadata-only synthetic evidence references.playbook-simulation.jsonwith tabletop decisions and no executed actions.decision-log.jsonlwith metadata-only response decisions.verification.jsonwith no-action/no-scan/no-PoC checks.manifest.jsonrecordingactive_scanning=false,external_network=false,exploit_poc=false,writes_to_host=false, andread_only=true.
P3 read-only reasoning artifacts land in the directory passed to
python -m phantom_secops.reasoning_scenario --out <dir>:
reasoning-report.jsonwith finding/evidence counts and read-only readiness.kill-chain-hypotheses.jsonwith synthetic advice-only hypotheses.playbook-review.jsonwith tabletop decision review and no executed actions.audit-summary.jsonwith metadata-only audit events.manifest.jsonrecordingactive_scanning=false,external_network=false,exploit_poc=false,writes_to_host=false,read_only=true, andactions_executed=false.
Full P3 contract: docs/REASONING_SCENARIO.md.
Alpha (0.1.0a0) and actively developed — but already a genuinely built,
comprehensively tested project rather than a scaffold. 400 tests pass, every
one behind an injectable command runner, so the suite exercises the real logic
while performing zero real scanning in CI (.github/workflows/ci.yml).
Shipped and tested:
- Red/blue kill-chain pipeline — 8-stage recon → vuln-scan → prose-only
exploit-suggest → report, with the honest, un-clamped MTTD metric, a
degradation "honesty gate" (a missing scanner shows
DEGRADED, never a fake clean run), and both a direct and a phantom-mesh driver. - Local-first endpoint self-check — host-posture audit (Windows + macOS),
dependency / OS-package CVE scanning (Trivy), and host intrusion detection
(a small Sigma engine over Windows event logs), merged by a deterministic,
no-LLM fusion step (
posture_fusion) into one prioritised action list. - Governed MCP orchestration — role-based deny policy, a fail-closed approval
gate with an operator CLI, and an append-only governance audit journal; every
tool carries
x-phantom.*capability metadata. - Offline MCP security scanner (
mcp_audit) — static checks for OWASP-MCP-Top-10-style risks (tool poisoning, SSRF, lethal-trifecta); it never connects to the servers it audits. phantom-secopsCLI —compare(diff two kill-chain reports/runs) plus the hermeticdefensive-loop,evidence-playbook, andreasoning-scenariodemo bundles.
Runs as an MCP server for the phantom mesh. The kill-chain façade
(python -m secops_mcp.server) exposes four composite tools — recon,
vuln_scan, detect, respond — that a phantom-mesh agent calls in
sequence, each tagged with x-phantom.* classification/capability metadata for
policy enforcement. Individual read-only endpoint tool servers (host audit, vuln
scan, IDS, self-audit, recon, log) are exposed alongside it — see
Architecture.
Roadmap (future direction, not blockers): wire the optional LLM-authored report path (narratives are deterministic templates today); add gated live-smoke runs against the Docker lab and a real phantom-mesh provider (external-tool logic is currently exercised through injected runners); and broaden the recon/tooling wrapper set.
These boundaries are the product, not limitations:
- Read-only / plain-text output. The exploit-suggester emits prose only —
has_runnable_pocis permanentlyfalse. Endpoint tools are read-only and self-scoped. It advises; it never changes your system. - Governed MCP orchestration. Every tool is tagged with
x-phantom.{classification, capabilities, read_only}capability metadata (e.g.blue/read.host_posture/target.self_only) — the hook for the per-agent policy enforcer in phantom-mesh, paired with its governor + phone-approval plane. - LLM as explainer / triager only. Mirroring what actually works in production (Semgrep Assistant,
Corgea, Socket): deterministic engines find facts, the LLM ranks / dedupes / explains. The
deterministic core (
posture_fusion) contains no LLM — it's the trustworthy spine.
Why this niche? The open-source offensive-agent space is a crowded gold-rush toward fully autonomous exploitation (Strix, CAI, PentAGI). The blue-team + endpoint-hygiene + governed-MCP lane is far less crowded — and the leading generic security-MCP bundles are ungoverned and now archived. The gap is exactly governance + a capability/policy model + a read-only posture. See the full landscape analysis in docs/phantom-secops.md.
┌────────────────────────────────────────────────────────────────────┐
│ phantom-mesh runtime │
│ LLM provider routing · tool-calling loop · cost tracking · │
│ inter-agent message passing · governor + phone-approval plane │
└─────────────┬───────────────────────────────────┬──────────────────┘
┌─────▼────────┐ ┌───────▼──────┐
│ RED agents │ (TOML-defined) │ BLUE agents │
└─────┬────────┘ └───────┬──────┘
┌─────▼────────┐ ┌───────▼──────┐
│ tool wrappers│ injectable runner │ tool wrappers│
│ (Python) │ + x-phantom tags │ (Python) │
└─────┬────────┘ └───────┬──────┘
docker exec│ into attacker docker socket │ → log volume
┌─────▼──────────────────────────────────── ▼──────┐
│ secops-lab docker network │
│ juice-shop · dvwa · metasploitable (targets) │
│ attacker (nmap/nuclei) · log-collector │
└───────────────────────────────────────────────────┘
8 engine modules, each a pure module behind an injectable command runner (so OS-touching logic is unit-tested with canned output — zero real scanning in CI):
| Capability | Engine | MCP tool |
|---|---|---|
| Host posture (firewall / disk encryption / AV / UAC / ports / SIP) | native OS queries | secops_host_audit |
| Dependency / OS-package CVEs (prioritised, fixable-first) | Trivy | secops_vuln |
| Host intrusion detection (encoded PowerShell, download cradles, AMSI bypass…) | small Sigma engine over Windows event logs | secops_ids |
Config self-audit (phantom-mesh agents.toml hygiene) |
native | secops_self_audit |
| Lab recon / log-anomaly (Pillar-1 tools, also exposed) | nmap / pattern matcher | secops_recon, secops_log |
The deterministic spine posture_fusion.fuse_posture merges host_audit + vuln_scan + ids_scan
into a single ranked action list (normalized severity, highest-risk-first, stable tiebreak,
plain-language, no LLM), wired into the real checkup.ps1.
- 400 passing tests, all behind injectable runners — no real scanning in the test suite.
CI in
.github/workflows/ci.yml. - Honest degradation, never a false alarm. Checks that need admin return
unknown+ a "re-run as administrator" hint, not a falsefail. A missing scanner in a live run shows a DEGRADED banner instead of a clean-looking "0 findings". - Low false-positives over coverage. An IDS rule that flagged a signed Microsoft module manifest as a download-cradle was tightened to 0 noise on 800 events. I deliberately don't stack 300+ CIS checks — for a personal machine that's alert fatigue, not security.
- Feed-don't-rescan. Deterministic findings are handed to the agent once; letting it re-scan a large repo timed out and falsely reported "no findings" while the log had 864.
- Security hardening of the tooling itself.
eval()→ a safe-AST boolean evaluator, nmap shell-injection patched, nuclei lab-gate fixed from a substring bypass to exact-hostname matching.
| ⛔ Red line | Why |
|---|---|
| No runnable PoC / exploit | has_runnable_poc is always false; the suggester emits prose only. This line is the product. |
| No external scanning | Lab targets are localhost / Docker-overlay only; endpoint tools are read-only and self-scoped. |
| No auto-remediation | Every tool advises, never changes your system. Until a human-in-the-loop approval model exists, you act — by design. |
| No autonomy drift | The gravity of this field is "let it self-exploit." Every step that way erases the niche and adds legal surface. |
| No "autonomous 0-day" headline chasing | That's frontier-lab + heavy-compute territory, not a solo Apache project's goal. |
Full legal scoping: ETHICS.md. Engineering decisions: docs/phantom-secops.md.