Skip to content

feat: add evidence-based Codex autonomy controls - #18

Merged
SignalLayerLabs merged 3 commits into
mainfrom
codex/marginal-python-runtime-discovery
Aug 14, 2026
Merged

SignalLayerLabs merged 3 commits into
mainfrom
codex/marginal-python-runtime-discovery

Conversation

@SignalLayerLabs

Copy link
Copy Markdown
Owner

What changed

  • add canonical Decision Receipts, structured utility/progress signals, and a secure v3 hash-chained governance ledger with deterministic v2 migration
  • add contextual L0-L4 authority and trust evaluation with explicit blockers, hysteresis, decay, verified-root transition receipts, and capability ceilings
  • add privacy-safe Italian/English user-intent normalization and authenticated Codex control-plane bypass
  • add deferred-consent Codex Autopilot with exact safe-repeat gating, per-workload pending isolation, fail-open recovery/demotion, and actual-only avoided/recovery counters
  • bind promotion to verified v3 evidence payload ranges; generic shell, tests, search, writes, network, deploy, and unknown MCP actions remain non-enforcing
  • add truthful status, doctor, explain, and privacy inspect commands and compatible Codex aliases
  • rebuild the deterministic plugin runtime and provenance

Why

MARGINAL previously had a credible Shadow Mode prototype, but promotion evidence was mutable,
authority was context-poor, natural user commands were not normalized, and repeated-action
enforcement was too broad. This change makes enforcement earned, narrowly scoped, auditable, and
recoverable while keeping correctness ahead of compute savings.

User impact

Users can install once, grant deferred Autopilot consent, and interact naturally in Italian or
English. MARGINAL remains quiet by default, never persists raw prompts or prompt hashes, and only
gates a constrained local read family after exact successful no-progress evidence. Status and
doctor report configured versus effective authority and the precise remaining blockers.

Validation

  • pytest -q — 571 passed
  • ruff format --check .
  • ruff check .
  • mypy src/marginal
  • deterministic Codex plugin build/check
  • independent P0/P1 reviews for receipts/privacy, v3 ledger/TOCTOU, trust, intent, Autopilot, and diagnostics; final re-reviews clean

Security and privacy

  • promotion requires a non-empty verified v3 root and exact evidence range
  • descriptor-relative no-follow file operations prevent parent-symlink races
  • failure/unknown outcomes clear the affected repeat proof and fail open
  • control-plane bypass requires trusted resolved launcher paths and an exact command grammar
  • raw prompts, commands, outputs, source payloads, credentials, and token-saving estimates are not persisted

Evidence limits and follow-up

The historical three-task SWE-bench smoke remains non-causal: both lanes resolved 0/3 tasks, so its
24.93% token difference is not evidence that MARGINAL caused savings. This PR does not claim causal
token reduction. Counterfactual/regret calibration, expanded public benchmark provenance, and the
larger repeated canary remain separate follow-up gates before broader compute enforcement or a
public marketplace release.

@SignalLayerLabs
SignalLayerLabs marked this pull request as ready for review August 14, 2026 08:47
@SignalLayerLabs
SignalLayerLabs merged commit 5a6586e into main Aug 14, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants