Skip to content

eval: real-vault fingerprint guard alarms on the orchestrator's own vault writes (false ABORT-REPORT) #750

Description

@toejough

Observed

During run opus-1789002029-6f1b03 (2026-09-09), the fingerprint guard raised a false positive ABORT-REPORT:

ABORT-REPORT: operator's real vault fingerprint changed! before=(1762, 1789001019.0935855) after=(1790, 1789002152.0255291). A trial may have reached real memory.

The 28 new files were:

  • 14 route-evidence notes written by the controller's recall glances (probe.py's own route skill exercises)
  • 14 .vec.json sidecars generated by those recalls

No trial touched the real vault (per-trial ENGRAM_VAULT_PATH verified, no write traces in trial logs).

Repro

  1. Start any probe run that fingerprints the real vault before/after:

    • See dev/eval/cumulative/runbook_vs_skill/probe.py::_real_vault_fingerprint() (lines 204-215)
    • And dev/eval/isolation.py::vault_fingerprint() / assert_vault_unchanged() (lines 176-199)
  2. In another shell, write to the real vault while the run is active:

    engram learn fact "example fact"
    # OR
    engram activate --note <any>
  3. The run ends with ABORT-REPORT even though the trial was isolated.

Current implementations compared:

  • probe.py::_real_vault_fingerprint() hashes: (file count, newest mtime)
  • isolation.py::vault_fingerprint() hashes: (note count, sha256 of sorted note basenames)

Both trigger on the orchestrator's own writes into the real vault.

Expected

The guard distinguishes trial-side leaks from session-side (orchestrator) writes. Options:

  1. Compare against a manifest with declared exclusions (allow route-evidence notes + sidecars during runs).
  2. Detect leaks trial-side: any write under the real vault from a trial's process tree, or any ENGRAM_VAULT_PATH unset in a trial's environment.
  3. Document the operating rule: "no real-vault writes during a fingerprinted run" and require operator discipline until detection is hardened.

Context

The orchestrator's own recall glances (from probe.py's route skill exercises) generate vault writes that the fingerprint guard cannot distinguish from trial leaks. This is a known issue (vault note 940). The guard is valuable for catching actual leaks, but needs refinement to avoid false alarms during multi-step orchestrated runs.

The run logs are at: /Users/joe/repos/personal/engram/.claude/worktrees/runbook-vs-skill/dev/eval/cumulative/runbook_vs_skill/results/opus_run.log

Affected files

  • dev/eval/cumulative/runbook_vs_skill/probe.py (lines 204-215: fingerprint guard logic)
  • dev/eval/isolation.py (lines 176-199: fingerprint comparison)

Related

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    needs-triageMaintainer needs to evaluate this issue

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions