diff --git a/demos/killer-demo/RESULTS.md b/demos/killer-demo/RESULTS.md index 808b851..61246d7 100644 --- a/demos/killer-demo/RESULTS.md +++ b/demos/killer-demo/RESULTS.md @@ -1,14 +1,18 @@ -# MARGINAL Killer Demo +# MARGINAL Demo 001 -**Fund only the next action worth taking.** +**AI agents repeat work that changed nothing. MARGINAL catches it.** + +Observe first. Prove waste. Earn enforcement. > Deterministic functional demonstration using declared action-cost estimates; not provider telemetry, not a production benchmark, and not a claim about every agent workload. +This deterministic artifact demonstrates MARGINAL's compute-selection discipline. The no-progress repetition sequence shown in the HTML is an explicitly labeled runtime-pattern illustration, not provider telemetry and not an enforcement benchmark. + Scenario: **Fix a percentage-discount bug in a deterministic Python repository** -![Baseline versus MARGINAL](comparison.svg) +![Without MARGINAL versus MARGINAL](comparison.svg) -## The defect +## Deterministic allocation proof ```diff - return total - rate @@ -19,7 +23,7 @@ Verifier: `apply_discount(100.0, 0.20) == 80.0` Initial verifier: **FAIL** -| Metric | Baseline: run everything | MARGINAL | Savings | +| Metric | Without MARGINAL | MARGINAL | Observed demo delta | |---|---:|---:|---:| | Declared tokens | 72,800 | 4,300 | **94.09%** | | Calls | 9 | 3 | **66.67%** | @@ -31,38 +35,31 @@ Initial verifier: **FAIL** ### Diagnose -Funded: **inspect the failing assertion** — approved: marginal ROI 7.166 +Selected: **inspect the failing assertion** — approved: marginal ROI 7.166 -| Candidate | Declared tokens | Expected gain | Score | Decision | -|---|---:|---:|---:|---| -| inspect the failing assertion | 1,200 | 0.220 | 0.189 | FUNDED: approved: marginal ROI 7.166 | -| scan the entire repository | 9,000 | 0.050 | -0.179 | SKIPPED: rejected: marginal ROI 0.218 below 1.000 | -| ask two parallel reviewers | 14,000 | 0.040 | -0.369 | SKIPPED: rejected: marginal ROI 0.098 below 1.000 | +Declared cost: **1,200 tokens · $0.006**. Alternatives rejected: **2**. ### Fix -Funded: **apply the targeted one-line patch** — approved: marginal ROI 7.418 +Selected: **apply the targeted one-line patch** — approved: marginal ROI 7.418 -| Candidate | Declared tokens | Expected gain | Score | Decision | -|---|---:|---:|---:|---| -| apply the targeted one-line patch | 2,400 | 0.500 | 0.433 | FUNDED: approved: marginal ROI 7.418 | -| rewrite the complete pricing module | 12,000 | 0.200 | -0.156 | SKIPPED: rejected: marginal ROI 0.561 below 1.000 | -| ask a frontier model for an alternative patch | 18,000 | 0.150 | -0.501 | SKIPPED: rejected: marginal ROI 0.230 below 1.000 | +Declared cost: **2,400 tokens · $0.018**. Alternatives rejected: **2**. ### Verify -Funded: **run the targeted verifier** — approved: marginal ROI 21.394 +Selected: **run the targeted verifier** — approved: marginal ROI 21.394 + +Declared cost: **700 tokens · $0.002**. Alternatives rejected: **2**. -| Candidate | Declared tokens | Expected gain | Score | Decision | -|---|---:|---:|---:|---| -| run the targeted verifier | 700 | 0.350 | 0.334 | FUNDED: approved: marginal ROI 21.394 | -| run the full test suite | 4,500 | 0.080 | -0.025 | SKIPPED: rejected: marginal ROI 0.763 below 1.000 | -| request a premium model audit | 11,000 | 0.050 | -0.348 | SKIPPED: rejected: marginal ROI 0.126 below 1.000 | +## What this demo proves + +- The deterministic task starts in FAIL and both workflows finish in PASS. +- The allocator can reject higher-cost actions while preserving the verifier outcome. +- Costs are declared demo estimates, not provider billing or production telemetry. +- The artifact is a mechanism demonstration, not a production benchmark. ## Reproduce ```bash marginal killer-demo --output killer-demo-output ``` - -The command writes this report, a standalone HTML report, an SVG comparison, the JSON result, and the provider-neutral decision trace. diff --git a/demos/killer-demo/comparison.svg b/demos/killer-demo/comparison.svg index f069068..1b60c69 100644 --- a/demos/killer-demo/comparison.svg +++ b/demos/killer-demo/comparison.svg @@ -1,25 +1,16 @@ - - - -Declared token cost for the same verified fix - Baseline - - -72,800 - MARGINAL - - -4,300 - -94.09% fewer declared tokens · outcome preserved + + + + +MARGINAL DEMO 001 +Same verified fix. Less declared demo compute. +Deterministic action-cost illustration — not provider telemetry. +Without MARGINAL · Baseline + +72,800 +With MARGINAL + +4,300 +94.09% fewer declared tokens +PASS → PASS · deterministic mechanism demonstration diff --git a/demos/killer-demo/index.html b/demos/killer-demo/index.html index 08e6b15..f060aec 100644 --- a/demos/killer-demo/index.html +++ b/demos/killer-demo/index.html @@ -2,1652 +2,52 @@ - - + + - MARGINAL Killer Demo - - - + MARGINAL Demo 001 — Stop AI Agent No-Progress Loops + + + + + + + + + + + + + + - -
-
- - - - -
-
-
-
Deterministic evaluation
-

Killer Demo

-

- Fund only the next action worth taking. - Same verified outcome. Far fewer tokens, lower cost, lower latency. -

- -
- - - - - Deterministic functional demonstration using declared action-cost estimates; not provider telemetry, not a production benchmark, and not a claim about every agent workload. -
-
- -
-
- Dante, the SignalLayer Labs mascot, representing MARGINAL compute allocation -
- - VERIFIED · PASS -
-
-
-
- -
-
-
-
- - - - - Token reduction -
- 94.09% - - 72,800 → 4,300 declared tokens - - -
- -
-
- - - - Actions executed -
- 9 → 3 - -
- -
-
- - - - Estimated cost -
- $0.763 → $0.026 - -
- -
-
- - - - - Estimated latency -
- 22.03s → 1.23s - -
- -
-
- - - - - Verified result -
- PASS → PASS - -
-
-
- -
-
-
-

Same fix. Different capital discipline.

-

- A like-for-like execution of diagnose, fix, and verification against the - same deterministic defect. -

-
-
- -
-
-
-
-
-

Baseline

- 9 actions -
-
  1. 1inspect the failing assertionDiagnose · 1,200 tokens$0.006Justified✓
  2. 2scan the entire repositoryDiagnose · 9,000 tokens$0.045Executedx
  3. 3ask two parallel reviewersDiagnose · 14,000 tokens$0.120Executedx
  4. 4apply the targeted one-line patchFix · 2,400 tokens$0.018Justified✓
  5. 5rewrite the complete pricing moduleFix · 12,000 tokens$0.110Executedx
  6. 6ask a frontier model for an alternative patchFix · 18,000 tokens$0.280Executedx
  7. 7run the targeted verifierVerify · 700 tokens$0.002Justified✓
  8. 8run the full test suiteVerify · 4,500 tokens$0.012Executedx
  9. 9request a premium model auditVerify · 11,000 tokens$0.170Executedx
-
- Total estimated cost - $0.763 -
-
- -
-
-

MARGINAL

- 3 actions -
-
  1. 1inspect the failing assertionDiagnose · 1,200 tokens$0.006Funded✓
  2. 2apply the targeted one-line patchFix · 2,400 tokens$0.018Funded✓
  3. 3run the targeted verifierVerify · 700 tokens$0.002Funded✓
-
- Total estimated cost - $0.026 -
-
-
-

- The baseline executes every available action. MARGINAL finances only the - economically justified sequence. -

-
- -
-
-

Allocation decisions

-

Every candidate is priced against expected marginal gain before execution.

-
-
- - - - - - - - - - -
ActionCostExpected gainDecision
inspect the failing assertionDiagnose$0.0061,200 tokens0.220score 0.189FUNDEDapproved: marginal ROI 7.166
scan the entire repositoryDiagnose$0.0459,000 tokens0.050score -0.179REJECTEDrejected: marginal ROI 0.218 below 1.000
ask two parallel reviewersDiagnose$0.12014,000 tokens0.040score -0.369REJECTEDrejected: marginal ROI 0.098 below 1.000
apply the targeted one-line patchFix$0.0182,400 tokens0.500score 0.433FUNDEDapproved: marginal ROI 7.418
rewrite the complete pricing moduleFix$0.11012,000 tokens0.200score -0.156REJECTEDrejected: marginal ROI 0.561 below 1.000
ask a frontier model for an alternative patchFix$0.28018,000 tokens0.150score -0.501REJECTEDrejected: marginal ROI 0.230 below 1.000
run the targeted verifierVerify$0.002700 tokens0.350score 0.334FUNDEDapproved: marginal ROI 21.394
run the full test suiteVerify$0.0124,500 tokens0.080score -0.025REJECTEDrejected: marginal ROI 0.763 below 1.000
request a premium model auditVerify$0.17011,000 tokens0.050score -0.348REJECTEDrejected: marginal ROI 0.126 below 1.000
-
-
-
-
- -
-
-
- -
-

What this demo proves

-
    -
  • The task starts in FAIL.
  • -
  • Both workflows finish in PASS.
  • -
  • Costs are declared action budgets used by the allocator.
  • -
  • This is a deterministic demonstration of economic action selection.
  • -
-
-
- -
-

The deterministic defect

-

Fix a percentage-discount bug in a deterministic Python repository

-
- - return total - rate - + return total * (1 - rate) -
-
- Verifier - apply_discount(100.0, 0.20) == 80.0 -
-
-
-
- -
- -
-

Build agents that spend compute deliberately.

-

- Explore the open-source project and run the same deterministic evaluation - locally. -

-
- -
- -
-
-

Reproduce the result

- -
- marginal killer-demo --output killer-demo-output -
- - -
-
+ +
+ +
+ +
+
+
MARGINAL · DEMO 001 · RUNTIME GOVERNOR

AI agents repeat work that changed nothing.

MARGINAL catches it. It observes no-progress repetition first, asks whether evidence actually changed, and only earns narrow authority to stop eligible repeats after the proof is strong enough.

Observe first. Prove waste. Earn enforcement.

+
Illustrative runtime pattern● LIVE TRACE
01
Read README.mdnew evidence acquired
RUN
02
Read README.mdverification pass
RUN
03
Read README.mdsame observable state
OBSERVE
04
Read README.mdsame action · same state · no new evidence
STOP CANDIDATE

Illustrative product mechanism — not provider telemetry and not this deterministic allocation benchmark.

+
+
The problem in five seconds

Activity is not progress.

A repeat is not automatically waste. MARGINAL looks for the stronger pattern: the same semantic action, unchanged observable state, and no new evidence.

WITHOUT MARGINALACTIVITY CONTINUES
01 · Read README.mdRUN
02 · Read README.mdRUN
03 · Read README.mdRUN AGAIN
04 · Read README.mdRUN AGAIN
05 · Read README.mdRUN AGAIN
WITH MARGINALEVIDENCE CHANGES THE DECISION
01 · New evidenceRUN
02 · VerificationRUN
03 · Same stateOBSERVE
04 · No-progress repeatSTOP CANDIDATE
+
Why it is different

Installing MARGINAL does not give it permission to block your agent.

Authority is evidence-backed and contextual. Shadow Mode can recommend a stop before MARGINAL is allowed to enforce one.

Shadow Mode · default

Recommendation without control.

MARGINAL can identify a stop candidate while the actual tool action still proceeds.

RecommendedSTOP
Actual behaviorALLOW
Earned Enforcement

Control has to be earned.

Only compatible, reviewed evidence can promote narrow enforcement. Drift, ambiguity, unknown outcomes, or safety failures demote authority and fail open.

AuthorityEARNED
ScopeNARROW
+
Decision path

Same action. Same state. No new evidence.

01Semantic repeat?Is the agent effectively attempting the same action again?
02State unchanged?Did the observable workspace stay the same?
03No new evidence?Did the previous pass fail to add useful evidence?
04Authority earned?If yes, an eligible repeat can become a stop candidate.
+
Deterministic mechanism proof

The real numbers in this artifact start here.

This section is generated from MARGINAL's deterministic allocator. It demonstrates action-selection economics, not Codex telemetry and not a no-progress enforcement benchmark.

Scope:Deterministic functional demonstration using declared action-cost estimates; not provider telemetry, not a production benchmark, and not a claim about every agent workload.
Token reduction94.09%
Declared tokens72,800 → 4,300
Actions9 → 3
Estimated USD$0.763 → $0.026
Verified resultPASS → PASS

Allocation decisions

Diagnoseinspect the failing assertion2 higher-cost alternatives rejected
Fixapply the targeted one-line patch2 higher-cost alternatives rejected
Verifyrun the targeted verifier2 higher-cost alternatives rejected

What this demo proves

Both deterministic workflows start from the same failing defect and end at the same verifier result. MARGINAL selects the targeted diagnose, fix, and verification actions using declared cost estimates. This does not establish production savings.

+
Fail-open by design

The governor should disappear when evidence is weak.

Changed stateRepeat pressure resets.
New evidenceThe next action is allowed.
Failure or unknownEnforcement fails open.
User requests repeatExplicit intent is respected.
+

See the code. Break the claim. Star it if it survives.

Open source, local first, provider neutral. The deterministic demo is reproducible with one command.

marginal killer-demo --output killer-demo-output
+
MARGINAL Killer Demo · Same verified outcome. Far fewer tokens, lower cost, lower latency. · Build agents that spend compute deliberately. · marginal-project-mark.png
+
+
- - +
+ - + \ No newline at end of file diff --git a/plugins/marginal/runtime/marginal_runtime.pyz b/plugins/marginal/runtime/marginal_runtime.pyz index c93eb5f..49b6fb9 100644 Binary files a/plugins/marginal/runtime/marginal_runtime.pyz and b/plugins/marginal/runtime/marginal_runtime.pyz differ diff --git a/plugins/marginal/runtime/provenance.json b/plugins/marginal/runtime/provenance.json index 6fafac6..0e6e74b 100644 --- a/plugins/marginal/runtime/provenance.json +++ b/plugins/marginal/runtime/provenance.json @@ -1 +1 @@ -{"builder":"scripts/build_codex_plugin.py","python_requires":">=3.10","schema_version":1,"sha256":"30c3268e4d1c051917f74f4b92ca60e1edd27f762096affcbb8b06fa7846271a","source_hash":"129c4898ab0da07d88b941adcb3c57d38ce69eb8a648ba5f4878266036855fba"} +{"builder":"scripts/build_codex_plugin.py","python_requires":">=3.10","schema_version":1,"sha256":"fca44c66cd52b7b79111f87288a84967d04cb0417fa2b7e55f6b22bfb41b3ce4","source_hash":"128c8e47bf46cb418e36a336d7a5754babec75c87704e577ec68a90acac450f3"} diff --git a/pyproject.toml b/pyproject.toml index 17e9257..4235086 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -98,6 +98,9 @@ src = ["src", "tests", "examples"] [tool.ruff.lint] select = ["E", "F", "I", "UP", "B", "SIM", "RUF"] +[tool.ruff.lint.per-file-ignores] +"src/marginal/killer_demo.py" = ["E501"] + [tool.ruff.format] quote-style = "double" indent-style = "space" diff --git a/src/marginal/killer_demo.py b/src/marginal/killer_demo.py index d67fd1b..e5641b2 100644 --- a/src/marginal/killer_demo.py +++ b/src/marginal/killer_demo.py @@ -424,17 +424,23 @@ def render_killer_demo_markdown(result: dict[str, Any]) -> str: marginal = result["marginal"] savings = result["savings"] lines = [ - "# MARGINAL Killer Demo", + "# MARGINAL Demo 001", "", - "**Fund only the next action worth taking.**", + "**AI agents repeat work that changed nothing. MARGINAL catches it.**", + "", + "Observe first. Prove waste. Earn enforcement.", "", f"> {result['disclaimer']}", "", + "This deterministic artifact demonstrates MARGINAL's compute-selection discipline. " + "The no-progress repetition sequence shown in the HTML is an explicitly labeled " + "runtime-pattern illustration, not provider telemetry and not an enforcement benchmark.", + "", f"Scenario: **{result['scenario']}**", "", - "![Baseline versus MARGINAL](comparison.svg)", + "![Without MARGINAL versus MARGINAL](comparison.svg)", "", - "## The defect", + "## Deterministic allocation proof", "", "```diff", f"- {result['defect']['before']}", @@ -445,7 +451,7 @@ def render_killer_demo_markdown(result: dict[str, Any]) -> str: "", "Initial verifier: **FAIL**", "", - "| Metric | Baseline: run everything | MARGINAL | Savings |", + "| Metric | Without MARGINAL | MARGINAL | Observed demo delta |", "|---|---:|---:|---:|", ( f"| Declared tokens | {baseline['tokens']:,} | {marginal['tokens']:,} | " @@ -469,35 +475,36 @@ def render_killer_demo_markdown(result: dict[str, Any]) -> str: "", ] for stage in result["stages"]: + funded = next( + candidate for candidate in stage["candidates"] if candidate["name"] == stage["selected"] + ) + rejected = len(stage["candidates"]) - 1 lines.extend( [ f"### {stage['stage']}", "", - f"Funded: **{stage['selected']}** — {stage['decision']}", + f"Selected: **{stage['selected']}** — {stage['decision']}", + "", + f"Declared cost: **{funded['tokens']:,} tokens · ${funded['usd']:.3f}**. " + f"Alternatives rejected: **{rejected}**.", "", - "| Candidate | Declared tokens | Expected gain | Score | Decision |", - "|---|---:|---:|---:|---|", ] ) - for candidate in stage["candidates"]: - status = "FUNDED" if candidate["name"] == stage["selected"] else "SKIPPED" - lines.append( - f"| {candidate['name']} | {candidate['tokens']:,} | " - f"{candidate['expected_gain']:.3f} | {candidate['score']:.3f} | " - f"{status}: {candidate['reason']} |" - ) - lines.append("") lines.extend( [ + "## What this demo proves", + "", + "- The deterministic task starts in FAIL and both workflows finish in PASS.", + "- The allocator can reject higher-cost actions while preserving the verifier outcome.", + "- Costs are declared demo estimates, not provider billing or production telemetry.", + "- The artifact is a mechanism demonstration, not a production benchmark.", + "", "## Reproduce", "", "```bash", "marginal killer-demo --output killer-demo-output", "```", "", - "The command writes this report, a standalone HTML report, an SVG comparison, " - "the JSON result, and the provider-neutral decision trace.", - "", ] ) return "\n".join(lines) @@ -506,35 +513,26 @@ def render_killer_demo_markdown(result: dict[str, Any]) -> str: def render_killer_demo_svg(result: dict[str, Any]) -> str: baseline_tokens = int(result["baseline"]["tokens"]) marginal_tokens = int(result["marginal"]["tokens"]) - chart_width = 520 - marginal_bar = max(4, round(chart_width * marginal_tokens / baseline_tokens)) savings = float(result["savings"]["tokens_percent"]) + max_width = 760 + marginal_width = max(8, round(max_width * marginal_tokens / baseline_tokens)) return "\n".join( [ - '', - ' ', - ' ', - "Declared token cost for the same verified fix", - ' Baseline', - f' ', - ' ', - f"{baseline_tokens:,}", - ' MARGINAL', - f' ', - f' ', - f"{marginal_tokens:,}", - ' ', - f"{savings:.2f}% fewer declared tokens · outcome preserved", + '', + '', + '', + '', + 'MARGINAL DEMO 001', + 'Same verified fix. Less declared demo compute.', + 'Deterministic action-cost illustration — not provider telemetry.', + 'Without MARGINAL · Baseline', + f'', + f'{baseline_tokens:,}', + 'With MARGINAL', + f'', + f'{marginal_tokens:,}', + f'{savings:.2f}% fewer declared tokens', + 'PASS → PASS · deterministic mechanism demonstration', "", "", ] @@ -632,1697 +630,81 @@ def render_killer_demo_html(result: dict[str, Any]) -> str: baseline = result["baseline"] marginal = result["marginal"] savings = result["savings"] - candidates = _candidate_lookup(result) - selected_names = {stage["selected"] for stage in result["stages"]} - baseline_steps = _render_flow_steps( - result["baseline_actions"], - candidates, - selected_names, - marginal=False, - ) - marginal_steps = _render_flow_steps( - result["marginal_actions"], - candidates, - selected_names, - marginal=True, + stage_cards = "".join( + "".join( + [ + '
', + f'
{html.escape(stage["stage"])}', + f"{html.escape(stage['selected'])}", + f"{len(stage['candidates']) - 1} higher-cost alternatives rejected", + "", + ] + ) + for stage in result["stages"] ) - allocation_rows = _render_allocation_rows(result) page = """ - - + + - MARGINAL Killer Demo - - - + MARGINAL Demo 001 — Stop AI Agent No-Progress Loops + + + + + + + + + + + + + + - -
-
- - - - -
-
-
-
Deterministic evaluation
-

Killer Demo

-

- Fund only the next action worth taking. - Same verified outcome. Far fewer tokens, lower cost, lower latency. -

- -
- - - - - {{DISCLAIMER}} -
-
- -
-
- Dante, the SignalLayer Labs mascot, representing MARGINAL compute allocation -
- - VERIFIED · PASS -
-
-
-
- -
-
-
-
- - - - - Token reduction -
- {{TOKEN_SAVINGS}}% - - {{BASELINE_TOKENS}} → {{MARGINAL_TOKENS}} declared tokens - - -
- -
-
- - - - Actions executed -
- {{BASELINE_CALLS}} → {{MARGINAL_CALLS}} - -
- -
-
- - - - Estimated cost -
- ${{BASELINE_USD}} → ${{MARGINAL_USD}} - -
- -
-
- - - - - Estimated latency -
- {{BASELINE_LATENCY}}s → {{MARGINAL_LATENCY}}s - -
- -
-
- - - - - Verified result -
- PASS → PASS - -
-
-
- -
-
-
-

Same fix. Different capital discipline.

-

- A like-for-like execution of diagnose, fix, and verification against the - same deterministic defect. -

-
-
- -
-
-
-
-
-

Baseline

- {{BASELINE_CALLS}} actions -
-
    {{BASELINE_STEPS}}
-
- Total estimated cost - ${{BASELINE_USD}} -
-
- -
-
-

MARGINAL

- {{MARGINAL_CALLS}} actions -
-
    {{MARGINAL_STEPS}}
-
- Total estimated cost - ${{MARGINAL_USD}} -
-
-
-

- The baseline executes every available action. MARGINAL finances only the - economically justified sequence. -

-
- -
-
-

Allocation decisions

-

Every candidate is priced against expected marginal gain before execution.

-
-
- - - - - - - - - - {{ALLOCATION_ROWS}} -
ActionCostExpected gainDecision
-
-
-
-
- -
-
-
- -
-

What this demo proves

-
    -
  • The task starts in FAIL.
  • -
  • Both workflows finish in PASS.
  • -
  • Costs are declared action budgets used by the allocator.
  • -
  • This is a deterministic demonstration of economic action selection.
  • -
-
-
- -
-

The deterministic defect

-

{{SCENARIO}}

-
- - {{DEFECT_BEFORE}} - + {{DEFECT_AFTER}} -
-
- Verifier - {{VERIFIER}} -
-
-
-
- -
- -
-

Build agents that spend compute deliberately.

-

- Explore the open-source project and run the same deterministic evaluation - locally. -

-
- -
- -
-
-

Reproduce the result

- -
- marginal killer-demo --output killer-demo-output -
- - -
-
+ +
+ +
+ +
+
+
MARGINAL · DEMO 001 · RUNTIME GOVERNOR

AI agents repeat work that changed nothing.

MARGINAL catches it. It observes no-progress repetition first, asks whether evidence actually changed, and only earns narrow authority to stop eligible repeats after the proof is strong enough.

Observe first. Prove waste. Earn enforcement.

+
Illustrative runtime pattern● LIVE TRACE
01
Read README.mdnew evidence acquired
RUN
02
Read README.mdverification pass
RUN
03
Read README.mdsame observable state
OBSERVE
04
Read README.mdsame action · same state · no new evidence
STOP CANDIDATE

Illustrative product mechanism — not provider telemetry and not this deterministic allocation benchmark.

+
+
The problem in five seconds

Activity is not progress.

A repeat is not automatically waste. MARGINAL looks for the stronger pattern: the same semantic action, unchanged observable state, and no new evidence.

WITHOUT MARGINALACTIVITY CONTINUES
01 · Read README.mdRUN
02 · Read README.mdRUN
03 · Read README.mdRUN AGAIN
04 · Read README.mdRUN AGAIN
05 · Read README.mdRUN AGAIN
WITH MARGINALEVIDENCE CHANGES THE DECISION
01 · New evidenceRUN
02 · VerificationRUN
03 · Same stateOBSERVE
04 · No-progress repeatSTOP CANDIDATE
+
Why it is different

Installing MARGINAL does not give it permission to block your agent.

Authority is evidence-backed and contextual. Shadow Mode can recommend a stop before MARGINAL is allowed to enforce one.

Shadow Mode · default

Recommendation without control.

MARGINAL can identify a stop candidate while the actual tool action still proceeds.

RecommendedSTOP
Actual behaviorALLOW
Earned Enforcement

Control has to be earned.

Only compatible, reviewed evidence can promote narrow enforcement. Drift, ambiguity, unknown outcomes, or safety failures demote authority and fail open.

AuthorityEARNED
ScopeNARROW
+
Decision path

Same action. Same state. No new evidence.

01Semantic repeat?Is the agent effectively attempting the same action again?
02State unchanged?Did the observable workspace stay the same?
03No new evidence?Did the previous pass fail to add useful evidence?
04Authority earned?If yes, an eligible repeat can become a stop candidate.
+
Deterministic mechanism proof

The real numbers in this artifact start here.

This section is generated from MARGINAL's deterministic allocator. It demonstrates action-selection economics, not Codex telemetry and not a no-progress enforcement benchmark.

Scope:%%DISCLAIMER%%
Token reduction%%TOKEN_SAVINGS%%%
Declared tokens%%BASELINE_TOKENS%% → %%MARGINAL_TOKENS%%
Actions%%BASELINE_CALLS%% → %%MARGINAL_CALLS%%
Estimated USD$%%BASELINE_USD%% → $%%MARGINAL_USD%%
Verified resultPASS → PASS

Allocation decisions

%%STAGE_CARDS%%

What this demo proves

Both deterministic workflows start from the same failing defect and end at the same verifier result. MARGINAL selects the targeted diagnose, fix, and verification actions using declared cost estimates. This does not establish production savings.

+
Fail-open by design

The governor should disappear when evidence is weak.

Changed stateRepeat pressure resets.
New evidenceThe next action is allowed.
Failure or unknownEnforcement fails open.
User requests repeatExplicit intent is respected.
+

See the code. Break the claim. Star it if it survives.

Open source, local first, provider neutral. The deterministic demo is reproducible with one command.

marginal killer-demo --output killer-demo-output
+
MARGINAL Killer Demo · Same verified outcome. Far fewer tokens, lower cost, lower latency. · Build agents that spend compute deliberately. · marginal-project-mark.png
+
+
- - +
+ - -""" +""" replacements = { - "{{MASCOT_URL}}": ( - "https://raw.githubusercontent.com/SignalLayerLabs/Marginal/main/assets/" - "marginal-project-mark.png" - ), - "{{DISCLAIMER}}": html.escape(result["disclaimer"]), - "{{TOKEN_SAVINGS}}": f"{savings['tokens_percent']:.2f}", - "{{BASELINE_TOKENS}}": f"{baseline['tokens']:,}", - "{{MARGINAL_TOKENS}}": f"{marginal['tokens']:,}", - "{{BASELINE_CALLS}}": str(baseline["calls"]), - "{{MARGINAL_CALLS}}": str(marginal["calls"]), - "{{BASELINE_USD}}": f"{baseline['usd']:.3f}", - "{{MARGINAL_USD}}": f"{marginal['usd']:.3f}", - "{{BASELINE_LATENCY}}": f"{baseline['latency_ms'] / 1000:.2f}", - "{{MARGINAL_LATENCY}}": f"{marginal['latency_ms'] / 1000:.2f}", - "{{BASELINE_STEPS}}": baseline_steps, - "{{MARGINAL_STEPS}}": marginal_steps, - "{{ALLOCATION_ROWS}}": allocation_rows, - "{{SCENARIO}}": html.escape(result["scenario"]), - "{{DEFECT_BEFORE}}": html.escape(result["defect"]["before"]), - "{{DEFECT_AFTER}}": html.escape(result["defect"]["after"]), - "{{VERIFIER}}": html.escape(result["defect"]["verifier"]), + "%%DISCLAIMER%%": html.escape(result["disclaimer"]), + "%%TOKEN_SAVINGS%%": f"{savings['tokens_percent']:.2f}", + "%%BASELINE_TOKENS%%": f"{baseline['tokens']:,}", + "%%MARGINAL_TOKENS%%": f"{marginal['tokens']:,}", + "%%BASELINE_CALLS%%": str(baseline["calls"]), + "%%MARGINAL_CALLS%%": str(marginal["calls"]), + "%%BASELINE_USD%%": f"{baseline['usd']:.3f}", + "%%MARGINAL_USD%%": f"{marginal['usd']:.3f}", + "%%STAGE_CARDS%%": stage_cards, } for placeholder, value in replacements.items(): page = page.replace(placeholder, value)