Meta published the blueprint. We measured the building.
Independent black-box verification of Meta's published Secure VM architecture, with new measurements from inside the runtime cell (htch-runtime).
While Meta published their high-level Secure VM design (Sept 8, 2026), this repository provides first-of-its-kind empirical measurements and observed runtime behaviors from inside the workload cell:
- Synthetic DNS VIP Allocation: RFC 2544 benchmark supernet (
198.18.0.0/15) dynamically mapped to external hosts, returningNOERROReven for non-existent domains. - Egress CA Certificate Details: On-the-fly TLS inspection and re-signing verified via locally installed per-instance egress CA trust anchors.
- AF_VSOCK Hypervisor Channel: Confirmed connectivity to CID 2 (the hypervisor host) on port 512, identifying the narrow guest-host boundary.
- Transparent Compression Accounting: 2 GB of zeroes written to
/var/cache/apt/archivesconsumed only ~58 MB of physical storage, verifying Btrfscompress-force=zstd:3overmount behavior. - The 100-Binary Suite & Multi-Call Engine: Full static dissection of
/opt/hatch/bin/, discovering the 17-applethatch-multicallarchitecture (Inode 1062),spawndeBPF gatekeeping, and pure Rust 1.97.1 toolchain. - Observed Sentinel Interception: Empirical capture of real-time stream scrubbing and transcript excision when credentials were leaked during diagnostics, forcing an immediate out-of-band agent re-steering.
This repository provides an empirical, non-invasive verification of Meta Muse (built on the internal Hatch platform runtime). Rather than analyzing static tool definitions or prompt wrappers, this research verifies the architecture from within the live runtime cell across three tracks:
- The Sandbox & Hypervisor Confinement (Track 1: The Cage): How untrusted agent code runs with container root privileges while strictly isolated via user namespaces (
uid 0 -> 131072), Seccomp BPF filters, Btrfs zstd overmounts, RFC 2544 synthetic DNS (198.18.0.0/15), and transparent egress proxies. - The Binary Engine & IPC Machinery (Track 2: The Machinery): The full 100-binary suite, the 17-applet
hatch-multicallmemory deduplication architecture,spawndlifecycle and eBPF cgroup supervisor,hatch-execdsubprocess cgroups, and the 38+ Unix domain socket matrix. - The Cognitive Engine & Sentinel Security (Track 3: The Mind): How enterprise autonomous agents evolve across multi-frequency asynchronous cadences (hourly memory consolidation, overnight goal studying, and nightly self-healing "dreaming" passes), shell hook event throttling (
silent()vswake()), and out-of-band Sentinel egress screening.
This research builds on and cross-verifies earlier disclosures, binary teardowns, and investigative reporting:
- Meta, "Security and safety for AI agents: our approach with Muse" (research.meta.ai, Sept 8, 2026): The canonical architecture disclosure describing the Secure VM,
systemd-nspawnruntime cell, Sentinel as sole egress authority, and credential surrogation. What this repo adds: Empirical verification from inside the running cell, mapping theoretical boundaries to exact syscall filters, network routes, and storage quotas.
- Hatch Engine Binary Teardown (Gist by @simonpure): Static reverse-engineering of an isolated single
hatchbinary, identifying 38 Unix sockets and JARVIS codenames. What this repo adds: Simon Pure analyzed only one 332 MB executable. Our work maps the entire 100-binary ecosystem, thehatch-multicallarchitecture, thespawndsupervisor, and combines static binary extraction with live in-situ runtime verification from inside the running cell. - macOS Client Teardown (muse-endo-teardown by @barkleesanders): Teardown of the Endo desktop client, 51-command device catalog, and Noise-encrypted WebSocket connection to per-user cloud VMs (measured 2026-09-17). What this repo adds: Focus on the cloud guest and container internals rather than client-side IPC.
- Sandbox Export Incident (ai-tldr.dev report): 6.8 GB runtime filesystem export via Google Drive (HN front page, 204 pts; closed as N/A by Meta). What this repo adds: Non-invasive, in-situ systems diagnostics of the live running runtime without bulk exfiltration.
- Jahanzaib.ai Architecture Breakdown (jahanzaib.ai): Summary of Meta's post, highlighting
io_uringremoval, capability stripping (CAP_SYS_PTRACE,CAP_NET_ADMIN), and containerized credentials. What this repo adds: Direct syscall probe matrices verifyingio_uringdenial and exact capability masks. - Forkast News Sentinel Analysis (forkast.news): Deep dive into Sentinel's kernel-level design, eBPF taint tracking, and credential surrogation (Sept 24, 2026). What this repo adds: Empirical capture of in-flight Sentinel output scrubbing and transcript excision during active agent execution.
- Ranzware Host Hardware Profiling (ranzware.com): Reported host specs (AMD EPYC Turin 9D25, 2 cores, 8 GB RAM, Ubuntu 24.04, Linux 7.0).
What this repo adds: In-guest
/proc/cpuinfo, DMI, and KVM hypervisor flag confirmation matching these specs. - The Information / Business Standard Reporting (Business Standard): Investigative coverage of the "hard gate" authorization model and credential-vault design. What this repo adds: In-product observation of how the cognitive loop, shell hooks, and Sentinel interact with the hard gate in real-world agent execution.
├── part1-the-cage/ # Track 1: Systems & Confinement Whitepaper
│ ├── README.md # Full 17-section whitepaper (Inside the Sandbox)
│ ├── evidence/ # Raw empirical probe logs & audit gap closure reports
│ │ ├── audit-gap-closure-report.md
│ │ ├── sandbox-telemetry-audit.md
│ │ └── vm-boundary-probe-report.md
│ └── tools/ # Reproduction tooling
│ ├── tailscale_ssh_proxy.py # Port 3130 stdio CONNECT tunnel helper
│ └── README.md
├── part2-the-machinery/ # Track 2: Binary Architecture, Multicall & IPC Whitepaper
│ ├── README.md # 100-binary suite, multicall dispatch, spawnd & IPC matrix
│ └── data/ # Extracted JSON schemas, crate deps, metrics & help dumps
│ ├── applet_dispatch.json
│ ├── crates.json
│ ├── ebpf.json
│ ├── inventory_table.md
│ ├── metrics.json
│ ├── routes_by_binary.json
│ ├── sockets_by_binary_v2.json
│ ├── tls.json
│ └── help/
└── part3-the-mind/ # Track 3: Cognitive Engine & Sentinel Whitepaper
├── README.md # Multi-cadence loops, dreaming, and Sentinel security
└── references/ # Synthetic examples of architecture references
├── alignment_synthesis_sample.md
├── hatch_hook_runtime.sh # silent() vs wake() event throttling wrapper
├── inferred_goal_leads_sample.md
└── self_improvement.md
-
📁 Part 1: The Cage — Sandbox Confinement, Storage & Network Mediation
- Full 17-section systems whitepaper covering the Cloud Hypervisor / KVM guest,
systemd-nspawncell, syscall error profile, Btrfs zstd/dev/mapper/rvovermounts, RFC 2544 synthetic DNS VIP pool, and the Port 3130 Tailscale CONNECT proxy. - Raw Diagnostic Evidence Logs
- Standalone Reproduction Tools (including
tailscale_ssh_proxy.py)
- Full 17-section systems whitepaper covering the Cloud Hypervisor / KVM guest,
-
📁 Part 2: The Machinery — Binary Architecture, Multi-Call Sub-Agents & IPC Topography
- Complete systems whitepaper on the 100-binary
/opt/hatch/bin/suite, Inode 1062hatch-multicalldeduplication,spawndeBPF cgroup gates,hatch-execdsocket activation, toolchain verification, and the 38+ socket IPC matrix. - Extracted Datasets & Schemas (crate graph, sockets, route tables, eBPF rules, metrics)
- Complete systems whitepaper on the 100-binary
-
📁 Part 3: The Mind — Cognitive Cadences, State Loops & Sentinel Security
- Autonomous agent systems whitepaper deconstructing the multi-frequency cognitive daemons, nightly alignment synthesis, latent goal inference engine (
INFERRED_GOAL_LEADS), shell hook throttling (silent()vswake()), and Sentinel IPC architecture. - Architecture References (synthetic examples of cognitive engine schemas and hook runtime)
- Autonomous agent systems whitepaper deconstructing the multi-frequency cognitive daemons, nightly alignment synthesis, latent goal inference engine (
sys-dissect is an independent systems research initiative dedicated to the empirical dissection, boundary verification, and containment analysis of production AI runtimes and autonomous agent infrastructure.
- Measure Runtime Realities: Architecture blueprints and design disclosures present intended models; empirical probing verifies actual kernel filters, hypervisor mediation, and storage behavior.
- Zero-Trust Boundary Analysis: Model token streams and prompt wrappers are never trusted for authorization; isolation guarantees must be physically enforced out-of-band by the operating system, container runtime, and hypervisor.
- Responsible & Non-Invasive: All diagnostics are conducted via bounded, black-box inspection inside authorized workload sessions. No exploits are staged, no host boundaries are breached, and all proprietary tokens or user telemetry are strictly sanitized before publication.
All findings were obtained via non-invasive, black-box diagnostic probing from within an authorized container session. No exploits were attempted, no host boundaries were breached, and no external services were disrupted.
All proprietary session identifiers, internal hostnames (*.metaaivm.com), access tokens, and personal user context have been rigorously scrubbed and de-identified.
Published by sys-dissect. Released under the MIT License.