feat: add evidence-based Codex autonomy controls - #18
Merged
Merged
Conversation
SignalLayerLabs
marked this pull request as ready for review
August 14, 2026 08:47
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed
status,doctor,explain, andprivacy inspectcommands and compatible Codex aliasesWhy
MARGINAL previously had a credible Shadow Mode prototype, but promotion evidence was mutable,
authority was context-poor, natural user commands were not normalized, and repeated-action
enforcement was too broad. This change makes enforcement earned, narrowly scoped, auditable, and
recoverable while keeping correctness ahead of compute savings.
User impact
Users can install once, grant deferred Autopilot consent, and interact naturally in Italian or
English. MARGINAL remains quiet by default, never persists raw prompts or prompt hashes, and only
gates a constrained local read family after exact successful no-progress evidence. Status and
doctor report configured versus effective authority and the precise remaining blockers.
Validation
pytest -q— 571 passedruff format --check .ruff check .mypy src/marginalSecurity and privacy
Evidence limits and follow-up
The historical three-task SWE-bench smoke remains non-causal: both lanes resolved 0/3 tasks, so its
24.93% token difference is not evidence that MARGINAL caused savings. This PR does not claim causal
token reduction. Counterfactual/regret calibration, expanded public benchmark provenance, and the
larger repeated canary remain separate follow-up gates before broader compute enforcement or a
public marketplace release.