This repository was archived by the owner on Aug 8, 2026. It is now read-only.
Close the forced-flow mechanism handoff: supported as written, marginal once the label is removed - #79
Merged
Merged
Conversation
…al once the label is removed Executed 2026-08-06 in npf, per section 7. Section 8 appended, body untouched. Verdict and every number live in npf/docs/crowdmon/2026-08-06-forced-flow-mechanism-verdict.md with its reproducer; this file carries a pointer and a one-line result, never a copy. Section 5.6's criteria return `supported` on both report types. That must not be quoted on its own: section 5.3's own placebo carries 73.8% (disaggregated) and 51.9% (tff) of the effect and is itself significant, where section 5.3 says it should be near zero if the crossing is what matters rather than the group label. The label IS the sign of pool_net, and pool_net mean-reverts, so a long book shrinking and a short book growing produces the predicted sign on weeks when nothing was triggered. Removing the label leaves -0.001388 (p 0.047) on disaggregated and -0.002657 (p 0.0099) on tff, about 1.2% and 1.9% of the pool against headline figures of 4.5% and 4.0%. Disaggregated would not survive a correction across even the two report types. A marginal lean, not a clean confirmation. Four places this specification was wrong or unexecutable, recorded in section 8 rather than corrected in the body: - Section 5.6 attaches no criterion to the placebo, so the criteria can be, and here are, satisfied by an effect that is mostly the group label. That gap is the most useful thing this run produced. - Section 4's control cannot be read from the panel and section 3 already says why: trigger_*_pool_agrees is 39 rows over one week. Recomputed over all 1,051 weeks from published pool_net plus propadj. pool_agrees reduces to the SIGN of pool_net and has no lookback dependence at all. - Section 4's "the two groups share the price mechanics exactly" is false as measured. The contradicted group crosses more often in both report types, 25.60% against 20.11% and 24.20% against 21.80%. - Section 5.5's p-value cannot be computed as written. Measured, p_literal runs 0.465 to 0.526 across all six variants including those whose recentred p_null is 0.0000. Section 3.1's equivalence reproduced exactly at 2.220446049250313e-16 over 96 comparisons rather than the required 12, all 96 signals matching. Store pinned: 267 parquets, zero moved during the run, byte copy outside the repo. Variant count 8, not 6, per section 5.8's own instruction. One declared deviation from section 6: the equivalence check runs in crowdmon's own venv rather than installing crowdmon into the shared npf/.venv. No number moves. Section 7's premise that npf has no register stopped being true on 2026-08-06. The file stays here anyway: it is tracked and now closed here, and moving a closed handoff would only cost a second lineage. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes
docs/handoffs/2026-08-05-forced-flow-mechanism-test.md, executed 2026-08-06 innpfby a session that has written none of this package. Body untouched, §8 appended, register row updated.Verdict, reproducer and artifacts:
npf/docs/crowdmon/2026-08-06-forced-flow-mechanism-verdict.md(npf PR #195). This file carries a pointer and a one-line result, never a copy.The answer
§5.6's criteria return
supportedon both report types, and that headline must not be quoted on its own.§5.3's own placebo carries 73.8% (Disaggregated) and 51.9% (TFF) of the effect, and is itself significant, where §5.3 says it "should be near zero if the crossing is what matters rather than the group label". The cause is mechanical: the label is the sign of
pool_net, andpool_netmean-reverts toward zero, so a long book shrinking and a short book growing produces the predicted sign on weeks when nothing was triggered.Removing the label leaves a real but small residual — a difference-in-differences of −0.001388 (p 0.047) and −0.002657 (p 0.0099), about 1.2% and 1.9% of the pool against headline figures of 4.5% and 4.0%. Disaggregated would not survive a correction across even the two report types. A marginal lean, not a clean confirmation.
Four places this specification was wrong or unexecutable
Recorded in §8 with the measurement that shows each, rather than corrected in the body.
trigger_*_pool_agreesis 39 rows over one week becauseadd_trigger_distanceis a point-in-time overlay. Recomputed over all 1,051 weeks from publishedpool_netpluspropadj. Worth carrying:pool_agreesreduces to the SIGN ofpool_netand has no lookback dependence at all.p_literalruns 0.465 to 0.526 across all six variants, including those whose recentredp_nullis 0.0000. Both reported; the criteria use the recentred one. Precedent for recording rather than silently applying: §7.7 of2026-08-02-validation-prereg.md.What held up
§3.1's equivalence check reproduced exactly:
2.220446049250313e-16over 96 comparisons rather than the required 12, all 96 signals matching. §5.6's asset-class condition passed at 9 of 10. Store pinned per §5.9: 267 parquets hashed, zero moved during the run, byte copy outside the repo.Stage C is descriptive by pre-registration and stays that way: realised flow sits near the bottom of the
[0, 2·|pool_net|]bracket, median position 0.090 and 0.083, which is a measurement of how thin the trend-following slice is in one week.Counts and deviations
8 variants, not §5.8's 6. The two additions are the difference-in-differences, one per report type — arithmetic on numbers the pre-registration already requires, made testable with its own bootstrap. Also added and carrying no statistic: a per-side descriptive table, and the equivalence check widened from 1 symbol to 8 (its figure is a maximum, so widening only tightens the gate).
One declared deviation from §6: the equivalence check runs in crowdmon's own venv rather than installing crowdmon into
npf/.venv, which is shared by the main checkout and every worktree. §6 is setup guidance, not a threshold, and no number moves. The phantom-package check §6 asks for is asserted in code before anything is computed.§7's premise has since changed and the file stays here anyway. §7 says it lives here because "
npfhas no equivalent" register;npfgrew one on 2026-08-06 and its companion moved there. This one is tracked and now closed here, and moving a closed handoff would only buy a second lineage of one document.Plain-language bottom line
When the rules say a trend-following book is about to be forced out at a price, and the price gets there, does the book actually shrink? Taken at face value, yes, and the scorecard says so. But most of that "yes" is an accounting artifact: the comparison group is defined by whether the book is long or short, and books drift back toward flat on their own. The pre-registration anticipated exactly this and built in a placebo, which was supposed to come back near zero and instead came back at half to three quarters of the headline.
Strip the artifact out and something real remains, in the predicted direction but small: roughly 1% to 2% of the book moving in the week after its trigger is hit. On the financial report that is reasonably solid; on the commodity report it is marginal enough that one fairness adjustment would erase it. Nothing here says the damage score predicts prices, and nothing licenses trading it.
🤖 Generated with Claude Code