Skip to content

feat(claude-code): Observe adapter and native plugin, with an engine-neutral hook core - #25

Merged
SignalLayerLabs merged 2 commits into
SignalLayerLabs:mainfrom
ralyodio:feat/hook-adapter-core-claude-code
Aug 17, 2026
Merged

SignalLayerLabs merged 2 commits into
SignalLayerLabs:mainfrom
ralyodio:feat/hook-adapter-core-claude-code

Conversation

@ralyodio

Copy link
Copy Markdown
Contributor

Problem

docs/integrations/overview.md lists Claude Code as roadmap work. Claude Code exposes a hook surface
that is close to the one the Codex plugin already uses, so an adapter is mostly a question of parsing
the right payloads and declaring honestly what that engine can prove.

Interface

Two additions:

marginal.integrations.hookkit — the parts of a hook integration that are genuinely engine
independent: normalized events (SessionBoundary, ToolCallStart, ToolCallEnd), privacy-safe action
normalization, conservative structured-outcome classification, workspace state evidence, and session
correlation (HookSessionRuntime). No economic policy: every decision comes from the core through
UniversalRuntime.

marginal.integrations.claude_code — the engine adapter: a strict parser for Claude Code's own
payloads, mapping onto hookkit, one authenticated loopback service per session, a reversible
installer that drives Claude Code's own plugin commands, and the plugin package under
plugins/marginal-claude-code/.

marginal install claude-code / marginal uninstall claude-code are wired into the existing CLI.

Behavior

Capability label: Observe. AgentCapabilities() with every flag false, so
UniversalRuntime refuses to run this adapter in a blocking mode. Every hook returns no output —
printing an advisory would change what the model sees, which is what Shadow Mode promises not to do —
so recommendations go to the Decision Ledger only.

Claude Code separates outcome at the event level: PostToolUse documents a call that succeeded and
PostToolUseFailure documents one that failed. Outcome is therefore an engine-declared fact, not an
inference from response text, and duration_ms gives measured latency that is settled as actual cost.
An interrupted call is recorded as unknown rather than charged as a tool failure. Per-tool token
usage is not exposed by the hook surface and is reported as unavailable rather than as zero.

Fail-open is unconditional. No plugin data directory, no reachable service, an unparsable payload, an
uninstallable marginal import: the hook exits 0 with no output and Claude Code proceeds as if the
plugin were absent.

Validation

Payload shapes were captured from Claude Code 2.1.233 rather than taken from documentation. The
captures showed two things worth recording: there is no model field in any hook payload (the Codex
parser requires one), and turn identity is prompt_id, not turn_id.

717 passed, 1 skipped (ruff format --check, ruff check, mypy src/marginal clean). 66 new tests
cover parsing, normalization, outcome classification, state observability, session correlation,
fail-open paths, the installer, and the shipped plugin manifests.

End-to-end against the real CLI: the plugin was installed with
claude plugin marketplace add + claude plugin install, then a session was asked to read one file
four times and run a failing shell command. The resulting ledger:

Event Recorded
read 1 APPROVED, NO_PROGRESS_UNOBSERVABLE
read 2, 3 APPROVED, NO_PROGRESS_OBSERVED
read 4 SHADOW_OVERRIDE, NO_PROGRESS_ENFORCEMENT_ELIGIBLE, recommended_stop: true, not blocked
failing shell outcome: failure from PostToolUseFailure, duration_ms: 349
session end 5 completed, 4 successful, 1 failed, 0 pending, 0 unmatched

Measured governance latency was 9.5–11.4 ms per decision. The ledger contained no file names, no file
contents, no command text, and no workspace paths.

Compatibility

marginal.integrations.codex.transport now re-exports the transport from
marginal.integrations.transport; the import path and behavior are unchanged and the existing Codex
tests pass untouched. The bundled Codex runtime zipapp and its provenance were rebuilt with
scripts/build_codex_plugin.py because the source tree changed.

Privacy

Tool arguments, command text, tool output, prompts, and transcripts are never persisted; only digests
are. transcript_path is dropped at parse time. Error text from a failed tool contributes to an
evidence digest and is not written to the ledger. The default profile is LOCAL_FULL because the
ledger stays in the plugin's own data directory on the user's machine; MARGINAL_PRIVACY_PROFILE
selects safe_telemetry when the ledger may cross a trust boundary, and there is a test asserting the
raw session identifier disappears under that profile. Pseudonymization is not anonymization.

Scientific limitations

  • Nothing here demonstrates that governing Claude Code saves tokens. This PR adds the measurement
    surface; a paired OFF/ON benchmark is separate work, and no performance claim is made.
  • The ~10 ms governance latency is a local single-machine observation from one session, not a
    characterized distribution.
  • Repetition detection assumes a Git work tree. Outside one the state hash is empty and every
    repetition control fails open, which means an unmeasured share of real sessions get no repetition
    signal at all.
  • Enforcement is deliberately absent. The documented deny transport is implemented and tested, but
    nothing calls it, because no Earned Enforcement evidence window exists for Claude Code. A stop
    recommendation stays a recommendation.
  • hookkit duplicates logic the Codex integration still carries privately. Migrating Codex onto the
    shared module is left as follow-up so a new adapter cannot destabilize the validated Codex path.

Adds a Claude Code integration labeled Observe, plus the engine-independent
parts of a hook integration extracted into `marginal.integrations.hookkit`.

The adapter records normalized tool-call evidence and repeated-work
recommendations in a local Decision Ledger. It declares no control capability,
so the core refuses to run it in a blocking mode, and every hook returns no
output: Shadow Mode never changes what Claude Code does next.

Claude Code reports success and failure as separate hook events
(`PostToolUse` and `PostToolUseFailure`), so outcomes are engine-declared facts
rather than inferences from response text, and `duration_ms` provides measured
latency. Per-tool token usage is not exposed and is reported as unavailable.

The authenticated loopback session transport moved out of the Codex package to
`marginal.integrations.transport`; the old import path re-exports it unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

@SignalLayerLabs SignalLayerLabs left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

CI is failing on UP036 for the Python version guard in the Claude hook bootstrap. I think the guard should stay because this script is invoked through the user's python3 and needs to fail open on an unsupported interpreter. Please suppress UP036 locally with a short comment explaining that, rather than removing the guard or disabling the rule globally. Then rerun the full CI matrix.

Comment thread plugins/marginal-claude-code/scripts/marginal_hook.py Outdated
@SignalLayerLabs
SignalLayerLabs merged commit 25b32eb into SignalLayerLabs:main Aug 17, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants