feat(claude-code): Observe adapter and native plugin, with an engine-neutral hook core - #25
Merged
SignalLayerLabs merged 2 commits intoAug 17, 2026
Conversation
Adds a Claude Code integration labeled Observe, plus the engine-independent parts of a hook integration extracted into `marginal.integrations.hookkit`. The adapter records normalized tool-call evidence and repeated-work recommendations in a local Decision Ledger. It declares no control capability, so the core refuses to run it in a blocking mode, and every hook returns no output: Shadow Mode never changes what Claude Code does next. Claude Code reports success and failure as separate hook events (`PostToolUse` and `PostToolUseFailure`), so outcomes are engine-declared facts rather than inferences from response text, and `duration_ms` provides measured latency. Per-tool token usage is not exposed and is reported as unavailable. The authenticated loopback session transport moved out of the Codex package to `marginal.integrations.transport`; the old import path re-exports it unchanged. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
SignalLayerLabs
left a comment
Owner
There was a problem hiding this comment.
CI is failing on UP036 for the Python version guard in the Claude hook bootstrap. I think the guard should stay because this script is invoked through the user's python3 and needs to fail open on an unsupported interpreter. Please suppress UP036 locally with a short comment explaining that, rather than removing the guard or disabling the rule globally. Then rerun the full CI matrix.
This was referenced Aug 17, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
docs/integrations/overview.mdlists Claude Code as roadmap work. Claude Code exposes a hook surfacethat is close to the one the Codex plugin already uses, so an adapter is mostly a question of parsing
the right payloads and declaring honestly what that engine can prove.
Interface
Two additions:
marginal.integrations.hookkit— the parts of a hook integration that are genuinely engineindependent: normalized events (
SessionBoundary,ToolCallStart,ToolCallEnd), privacy-safe actionnormalization, conservative structured-outcome classification, workspace state evidence, and session
correlation (
HookSessionRuntime). No economic policy: every decision comes from the core throughUniversalRuntime.marginal.integrations.claude_code— the engine adapter: a strict parser for Claude Code's ownpayloads, mapping onto
hookkit, one authenticated loopback service per session, a reversibleinstaller that drives Claude Code's own plugin commands, and the plugin package under
plugins/marginal-claude-code/.marginal install claude-code/marginal uninstall claude-codeare wired into the existing CLI.Behavior
Capability label: Observe.
AgentCapabilities()with every flag false, soUniversalRuntimerefuses to run this adapter in a blocking mode. Every hook returns no output —printing an advisory would change what the model sees, which is what Shadow Mode promises not to do —
so recommendations go to the Decision Ledger only.
Claude Code separates outcome at the event level:
PostToolUsedocuments a call that succeeded andPostToolUseFailuredocuments one that failed. Outcome is therefore an engine-declared fact, not aninference from response text, and
duration_msgives measured latency that is settled as actual cost.An interrupted call is recorded as
unknownrather than charged as a tool failure. Per-tool tokenusage is not exposed by the hook surface and is reported as unavailable rather than as zero.
Fail-open is unconditional. No plugin data directory, no reachable service, an unparsable payload, an
uninstallable
marginalimport: the hook exits 0 with no output and Claude Code proceeds as if theplugin were absent.
Validation
Payload shapes were captured from Claude Code 2.1.233 rather than taken from documentation. The
captures showed two things worth recording: there is no
modelfield in any hook payload (the Codexparser requires one), and turn identity is
prompt_id, notturn_id.717 passed, 1 skipped(ruff format --check,ruff check,mypy src/marginalclean). 66 new testscover parsing, normalization, outcome classification, state observability, session correlation,
fail-open paths, the installer, and the shipped plugin manifests.
End-to-end against the real CLI: the plugin was installed with
claude plugin marketplace add+claude plugin install, then a session was asked to read one filefour times and run a failing shell command. The resulting ledger:
APPROVED,NO_PROGRESS_UNOBSERVABLEAPPROVED,NO_PROGRESS_OBSERVEDSHADOW_OVERRIDE,NO_PROGRESS_ENFORCEMENT_ELIGIBLE,recommended_stop: true, not blockedoutcome: failurefromPostToolUseFailure,duration_ms: 349Measured governance latency was 9.5–11.4 ms per decision. The ledger contained no file names, no file
contents, no command text, and no workspace paths.
Compatibility
marginal.integrations.codex.transportnow re-exports the transport frommarginal.integrations.transport; the import path and behavior are unchanged and the existing Codextests pass untouched. The bundled Codex runtime zipapp and its provenance were rebuilt with
scripts/build_codex_plugin.pybecause the source tree changed.Privacy
Tool arguments, command text, tool output, prompts, and transcripts are never persisted; only digests
are.
transcript_pathis dropped at parse time. Error text from a failed tool contributes to anevidence digest and is not written to the ledger. The default profile is
LOCAL_FULLbecause theledger stays in the plugin's own data directory on the user's machine;
MARGINAL_PRIVACY_PROFILEselects
safe_telemetrywhen the ledger may cross a trust boundary, and there is a test asserting theraw session identifier disappears under that profile. Pseudonymization is not anonymization.
Scientific limitations
surface; a paired OFF/ON benchmark is separate work, and no performance claim is made.
characterized distribution.
repetition control fails open, which means an unmeasured share of real sessions get no repetition
signal at all.
nothing calls it, because no Earned Enforcement evidence window exists for Claude Code. A stop
recommendation stays a recommendation.
hookkitduplicates logic the Codex integration still carries privately. Migrating Codex onto theshared module is left as follow-up so a new adapter cannot destabilize the validated Codex path.