M0 audit corrections (H1-H3 + M4-M7; NEW-1 disclosed) — do not merge pending re-review - #2
Conversation
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…calibration (H1,H2,H3,M4-parser) - H1: attribute the primary loss by lost contribution per channel (two-tier report); remove the 88%-util flip - H2: menu-positioning coherence order-weight (zero below 2.0x median; goldens byte-identical) - H3: forecast realization factor + comp haircut + honest downward-skewed band; demand-opportunity vs completed labels - M4: Money.TryParse (integer cents); M5: opt-in dequeue trace; preserve the discovered hybrid as a regression fixture Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…(M4) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…o the harness Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…nce, FIFO order, labor, price parse, widened no-float guard) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ate M0 evidence Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Independent Re-Review — M0 Correction PR #2A. Review identity
B. Executive verdict
C. Finding resolution table
D. Causal-attribution results (H1)The core defect is fixed. E. Forecast-validation results (H3) — held-outEvaluated on held-out seeds not used in development (33,000 runs):
F. Strategy-integrity results (H2 + NEW-1)
G. Medium findings (evidence)
H. Golden-checksum explanationThe three goldens are byte-identical to pre-fix and this is correct and expected: I. Scope resultCorrection-only, confirmed. Delta is prod +376/−60 (9 files), test +258/−5, docs/reports +280/−25. No executable repeat-visit / reputation / reviews / critics / marketing / inventory / suppliers / graphics / engine / pathfinding / save / multi-restaurant / manager / delegation / campaign / audio / asset system was added (the only "reputation" hits are disclosure strings in a report). Core stays pure (no File/Directory/async/Random). 103 tests pass twice. J. Human-test readinessNot ready. Two contaminants of "can the player form a plan and be fairly informed": (1) the forecast reliably points players the wrong way on seat count (the most basic capacity lever), and is badly biased for the plans a learning player builds; (2) pricing is degenerate — "overprice everything ~2.5×" dominates, so a perceptive tester experiences a shallow pricing decision, and a min-maxer trivially breaks the balance. The autopsy (H1) is now trustworthy, but the forecast is not, and pricing lacks a real tradeoff. Fix both in a second bounded correction, then run the five-player gate on a build where all four decision axes are fairly informed. K. Required corrections (for M0 only — no M1)
L. Final gateThe corrections are real and the build is honest and in scope, but the "no single dominant strategy" clause of the M0 product question fails (a robust universal overprice dominator exists), and the forecast is directionally misleading on capacity decisions. Both are fixable within M0 via a second small strategy-integrity + forecast correction. Not M. Required owner decisionHoward & Aaron: authorize a second bounded M0 correction (add a real single-service price-elasticity cost so overpricing is not universally optimal, and fix the forecast so excess seats lower expected contribution), then run the five-player human gate on the corrected build — or explicitly accept shipping M0 with a known robust universal dominator and a directionally-unreliable forecast (not recommended). "M0.5" is not an approved milestone; this is a second correction inside M0, not M1. The PR #2 branch was not modified during this review. Neither PR was merged. M1 was not begun. |
M0 bounded correction pass (responds to the PR #1 independent audit)
Branched from the reviewed head
beb9289. Scope-locked to the audit's 3 High + 5 Medium findings — no M1, no new features, no content expansion. Base = the foundation branch so this diff is exactly the audit response. Do not merge yet: awaiting independent re-review + owner-run human playtests.High
AttributionTests.DominanceRegressionTests.ForecastCalibrationTests.Medium
Money.TryParse, nodouble). M5 direct FIFO dequeue-order test + labor-value test. M6 no-float guard widened to static fields + properties. M7 cross-OS determinism claims corrected to "verified same-environment only."Honestly disclosed, NOT hidden (NEW-1)
The dominance search still finds one residual dominator: a uniformly-overpriced coherent menu. It stems from M0 having no repeat-visit/reputation teeth on price/quality (the disclosed throughput boundary, D-012). A price-elasticity fix would change a golden fixture, so it is deferred to M0.5, not hacked here. See
reports/balance/dominance-search.md.Verification
103 tests pass (was 74); goldens byte-identical; determinism holds; accounting reconciles; scope clean; no assets/binaries committed. Full write-up:
reports/m0/M0-CORRECTION-REPORT.md; pre-fix reproductions:reports/m0/PRE-FIX-AUDIT-EVIDENCE.md.Gate: Conditional-Pass / Defer. M1 remains unauthorized.
🤖 Generated with Claude Code