Skip to content

claude-code: research a defensible Tool Enforcement boundary #64

Description

@SignalLayerLabs

Question

Does Claude Code expose a sufficiently reliable interception boundary for MARGINAL to safely implement Tool Enforcement for any constrained action family?

Contribution wanted

Research and prototype the boundary without broadening production authority.

The goal is not to force Claude Code into parity with Codex. The goal is to establish, with reproducible evidence, whether a defensible enforcement boundary actually exists.

Required work

  • Identify the exact native Claude Code mechanism used for interception.
  • Record exact Claude Code version(s) tested.
  • Determine which action/tool families can actually be intercepted.
  • Determine whether denial happens before side effects.
  • Verify what happens on timeout, malformed hook output, plugin crash and unavailable runtime.
  • Characterize fail-open vs fail-closed behavior.
  • Determine whether concurrent/subagent execution changes the interception guarantee.
  • Determine whether internal, specialized or alternate tool paths bypass interception.
  • Characterize lifecycle coverage required to correlate proposal → deny/allow → outcome.
  • Document all known blind spots.

Capability discipline

A prompt instruction, advisory message, injected context or best-effort middleware is not Tool Enforcement.

Transporting a deny-like directive is also insufficient by itself.

To qualify as Tool Enforcement, the integration must demonstrate that MARGINAL can intercept a defined action family and prevent the action before its side effect occurs.

Do not generalize evidence from one action family to all Claude Code tools.

Safety requirements

Any prototype must:

  • fail open when MARGINAL is unavailable unless the platform guarantees a safer documented behavior;
  • avoid leaving Claude Code unusable after plugin/runtime failure;
  • preserve current Observe behavior by default;
  • not silently enable enforcement for existing users;
  • keep authority local and evidence-earned;
  • not use Commons/shared popularity as promotion evidence.

Privacy

Do not persist raw:

  • prompts;
  • source code;
  • command text;
  • tool output;
  • private paths;
  • repository identity;
  • credentials;
  • persistent user/device/install identifiers.

Use synthetic fixtures for adversarial interception tests.

Acceptance criteria

Deliver reproducible evidence supporting one of these outcomes:

Outcome A — defensible boundary exists

Document precisely:

  • the supported action family;
  • the tested Claude Code version;
  • the interception mechanism;
  • pre-side-effect denial proof;
  • known bypasses/blind spots;
  • failure behavior;
  • tests demonstrating the boundary.

This does not automatically enable enforcement. A separate Earned Enforcement change is still required.

Outcome B — boundary is insufficient

Document why Claude Code must remain Observe for the tested action family and what technical evidence is missing.

Both outcomes are valid.

Non-goal

Do not enable production Claude Code Tool Enforcement in this issue.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    claude-codeClaude Code integrationevidenceEvidence quality, attribution and validationgovernanceGovernance, authority and enforcement policyhelp wantedExtra attention is neededintegrationAgent/runtime integration workresearchEvidence gathering or falsification work

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions