Skip to content
This repository was archived by the owner on Aug 8, 2026. It is now read-only.

Close the forced-flow mechanism handoff: supported as written, marginal once the label is removed - #79

Merged
mspinola merged 1 commit into
mainfrom
claude/ffm-outcome
Aug 7, 2026
Merged

Close the forced-flow mechanism handoff: supported as written, marginal once the label is removed#79
mspinola merged 1 commit into
mainfrom
claude/ffm-outcome

Conversation

@mspinola

@mspinola mspinola commented Aug 6, 2026

Copy link
Copy Markdown
Owner

Closes docs/handoffs/2026-08-05-forced-flow-mechanism-test.md, executed 2026-08-06 in npf by a session that has written none of this package. Body untouched, §8 appended, register row updated.

Verdict, reproducer and artifacts: npf/docs/crowdmon/2026-08-06-forced-flow-mechanism-verdict.md (npf PR #195). This file carries a pointer and a one-line result, never a copy.

The answer

§5.6's criteria return supported on both report types, and that headline must not be quoted on its own.

§5.3's own placebo carries 73.8% (Disaggregated) and 51.9% (TFF) of the effect, and is itself significant, where §5.3 says it "should be near zero if the crossing is what matters rather than the group label". The cause is mechanical: the label is the sign of pool_net, and pool_net mean-reverts toward zero, so a long book shrinking and a short book growing produces the predicted sign on weeks when nothing was triggered.

Removing the label leaves a real but small residual — a difference-in-differences of −0.001388 (p 0.047) and −0.002657 (p 0.0099), about 1.2% and 1.9% of the pool against headline figures of 4.5% and 4.0%. Disaggregated would not survive a correction across even the two report types. A marginal lean, not a clean confirmation.

Four places this specification was wrong or unexecutable

Recorded in §8 with the measurement that shows each, rather than corrected in the body.

  • §5.6 attaches no criterion to the placebo, so the criteria can be — and here are — satisfied by an effect that is mostly the group label. This is the most useful thing the run produced, and anything re-running the test should fix it first.
  • §4's control cannot be read from the panel, and §3 already says why. trigger_*_pool_agrees is 39 rows over one week because add_trigger_distance is a point-in-time overlay. Recomputed over all 1,051 weeks from published pool_net plus propadj. Worth carrying: pool_agrees reduces to the SIGN of pool_net and has no lookback dependence at all.
  • §4's central design claim is false as measured. "The two groups share the price mechanics exactly": the contradicted group crosses more often in both report types, 25.60% against 20.11% and 24.20% against 21.80%, because agreement selects weeks whose trend is less likely to reverse inside five sessions.
  • §5.5's p-value cannot be computed as written. A bootstrap resampling the observed data is centred on the observed statistic. Measured, p_literal runs 0.465 to 0.526 across all six variants, including those whose recentred p_null is 0.0000. Both reported; the criteria use the recentred one. Precedent for recording rather than silently applying: §7.7 of 2026-08-02-validation-prereg.md.

What held up

§3.1's equivalence check reproduced exactly: 2.220446049250313e-16 over 96 comparisons rather than the required 12, all 96 signals matching. §5.6's asset-class condition passed at 9 of 10. Store pinned per §5.9: 267 parquets hashed, zero moved during the run, byte copy outside the repo.

Stage C is descriptive by pre-registration and stays that way: realised flow sits near the bottom of the [0, 2·|pool_net|] bracket, median position 0.090 and 0.083, which is a measurement of how thin the trend-following slice is in one week.

Counts and deviations

8 variants, not §5.8's 6. The two additions are the difference-in-differences, one per report type — arithmetic on numbers the pre-registration already requires, made testable with its own bootstrap. Also added and carrying no statistic: a per-side descriptive table, and the equivalence check widened from 1 symbol to 8 (its figure is a maximum, so widening only tightens the gate).

One declared deviation from §6: the equivalence check runs in crowdmon's own venv rather than installing crowdmon into npf/.venv, which is shared by the main checkout and every worktree. §6 is setup guidance, not a threshold, and no number moves. The phantom-package check §6 asks for is asserted in code before anything is computed.

§7's premise has since changed and the file stays here anyway. §7 says it lives here because "npf has no equivalent" register; npf grew one on 2026-08-06 and its companion moved there. This one is tracked and now closed here, and moving a closed handoff would only buy a second lineage of one document.

Plain-language bottom line

When the rules say a trend-following book is about to be forced out at a price, and the price gets there, does the book actually shrink? Taken at face value, yes, and the scorecard says so. But most of that "yes" is an accounting artifact: the comparison group is defined by whether the book is long or short, and books drift back toward flat on their own. The pre-registration anticipated exactly this and built in a placebo, which was supposed to come back near zero and instead came back at half to three quarters of the headline.

Strip the artifact out and something real remains, in the predicted direction but small: roughly 1% to 2% of the book moving in the week after its trigger is hit. On the financial report that is reasonably solid; on the commodity report it is marginal enough that one fairness adjustment would erase it. Nothing here says the damage score predicts prices, and nothing licenses trading it.

🤖 Generated with Claude Code

…al once the label is removed

Executed 2026-08-06 in npf, per section 7. Section 8 appended, body untouched.
Verdict and every number live in npf/docs/crowdmon/2026-08-06-forced-flow-mechanism-verdict.md
with its reproducer; this file carries a pointer and a one-line result, never a
copy.

Section 5.6's criteria return `supported` on both report types. That must not be
quoted on its own: section 5.3's own placebo carries 73.8% (disaggregated) and
51.9% (tff) of the effect and is itself significant, where section 5.3 says it
should be near zero if the crossing is what matters rather than the group label.
The label IS the sign of pool_net, and pool_net mean-reverts, so a long book
shrinking and a short book growing produces the predicted sign on weeks when
nothing was triggered.

Removing the label leaves -0.001388 (p 0.047) on disaggregated and -0.002657
(p 0.0099) on tff, about 1.2% and 1.9% of the pool against headline figures of
4.5% and 4.0%. Disaggregated would not survive a correction across even the two
report types. A marginal lean, not a clean confirmation.

Four places this specification was wrong or unexecutable, recorded in section 8
rather than corrected in the body:

- Section 5.6 attaches no criterion to the placebo, so the criteria can be, and
  here are, satisfied by an effect that is mostly the group label. That gap is the
  most useful thing this run produced.
- Section 4's control cannot be read from the panel and section 3 already says
  why: trigger_*_pool_agrees is 39 rows over one week. Recomputed over all 1,051
  weeks from published pool_net plus propadj. pool_agrees reduces to the SIGN of
  pool_net and has no lookback dependence at all.
- Section 4's "the two groups share the price mechanics exactly" is false as
  measured. The contradicted group crosses more often in both report types,
  25.60% against 20.11% and 24.20% against 21.80%.
- Section 5.5's p-value cannot be computed as written. Measured, p_literal runs
  0.465 to 0.526 across all six variants including those whose recentred p_null is
  0.0000.

Section 3.1's equivalence reproduced exactly at 2.220446049250313e-16 over 96
comparisons rather than the required 12, all 96 signals matching. Store pinned:
267 parquets, zero moved during the run, byte copy outside the repo.

Variant count 8, not 6, per section 5.8's own instruction. One declared deviation
from section 6: the equivalence check runs in crowdmon's own venv rather than
installing crowdmon into the shared npf/.venv. No number moves.

Section 7's premise that npf has no register stopped being true on 2026-08-06.
The file stays here anyway: it is tracked and now closed here, and moving a closed
handoff would only cost a second lineage.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@mspinola
mspinola merged commit 0f6169e into main Aug 7, 2026
5 checks passed
@mspinola
mspinola deleted the claude/ffm-outcome branch August 7, 2026 00:03
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant