M0 second correction: strategy integrity (price elasticity) + forecast direction — do not merge pending re-review - #3
Conversation
…rrection) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Willingness-to-pay is anchored to each dish's own calibrated suggested price (not the menu median a uniform-overprice menu hides behind), widened by segment tolerance and dish quality. Above WTP, order probability decays smoothly (hyperbolic), scaled by segment price sensitivity, floored at 3%. No hard cap, no name/recipe branch, no single-threshold cliff. Suggested prices recentered (~+40%) to give the model credible absolute anchors and segment sensitivities retuned (Value 8000->10000, Social 5000->5500, Enthusiast 3500->3000) so the segments span a clear elasticity range. Wired into demand conversion and dish choice. PriceResistScaleBp=450 from an elasticity target (max-sensitivity diner keeps ~1/3 at ~8% over WTP), not the dominance outcome. Adds the UniformlyOverpricedCoherentMenu regression fixture. See docs/design/PRICING-CONTRACT.md, DECISION-LOG D-020/D-022/D-023. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…forecast (H3) Kitchen cover capacity is now the per-station bottleneck (min over stations), realization is keyed off kitchen throughput only, and seats past the kitchen wall incur an over-acceptance waste cost (proportional to the seat surplus and demand pressure; slack demand wastes nothing). Together these stop the forecast reversing the seat decision (was 46->66 seats: forecast +$483 / actual -$377). See docs/design/FORECAST-CONTRACT.md, DECISION-LOG D-021. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Replaces the weak "does any plan beat the NAMED strategies" test (low seeds, under-optimized bar) with a searched per-market frontier (random + archetype sweep, 80 seeds/market). Reports whether any single plan is within $150 of the frontier in all three markets. Drove by an independent verifier that showed a plan can beat every named winner without being a cross-market optimum. See DECISION-LOG D-024. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Adds PriceElasticityTests (per-dish WTP anchoring, smoothness/no-cliff, overprice unprofitability, segment ordering, uniform-overprice menu beaten), ForecastDirectionTests (seats past the kitchen wall / needed cook / overpricing), and DominanceFrontierTests (opposed per-market regimes, no plan near-optimal everywhere, the searched champion is not a dominator). Updates the forecast calibration bounds and re-baselines the three golden checksums for the deliberate price/sensitivity recentering (DECISION-LOG D-022). 117 tests pass. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Adds PRICING-CONTRACT.md and FORECAST-CONTRACT.md; records D-020..D-024 (elasticity, forecast fix, golden re-baseline, contextual premium, frontier dominance methodology); updates balance hypotheses with the frontier result and disclosed M0.5 residuals (named-set enrichment, labor/throughput calibration); refreshes CURRENT-STATE and NEXT-ACTION. Corrects the earlier "0 dominators" wording to the honest frontier claim. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Regenerates distribution, dominance-search (frontier-based), forecast calibration, determinism, and example forecast at 200 seeds, and adds the §23 A-K builder's report (M0-CORRECTION-2-REPORT.md) recommending a technical Conditional-Pass pending owner-run human playtests. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Focused Independent Re-Review — M0 Pricing & Forecast Integrity (PR #3)A. Review identity
B. Executive verdict
The engine is genuinely sound and the two chartered fixes are partially real: the specific H2 uniform-overprice exploit is closed, the H3 seat-direction reversal is fixed, determinism/accounting/scope/purity all pass, and price elasticity is principled (single-peaked, smooth, per-dish, substitution/decline work). But two load-bearing claims are not earned and would contaminate the human verdict: (1) the forecast recommends the wrong pricing direction in the value-lunch market, and (2) the flagship "no cross-market dominator / structurally-opposed regimes" claim is refuted by stronger independent search — a single fixed premium plan is broadly near-optimal in every market and wins the value lunch. Both are correctable without an engine rewrite, so Defer (not Rewrite).
C. Finding table
D. Price-model evidence (Domain A — reproduced)
E. Strategy-frontier evidence (Domain B — reproduced, contradicts the central claim)Independent multi-restart hill-climbing (a different algorithm than the builder's random+archetype sweep) found a single fixed, legal plan —
F. Forecast evidence (Domain C — reproduced)
G. Labor evidence (Domain E — non-blocking)Understaffing is severely, legibly costly (a starved station ⇒ walkouts, 90–100% loss). A genuine lean strategy is viable with menu/station discipline (2 cooks + FOH on a 2-station menu: +$1186…+$1462, 0% loss). But on a fixed menu/seats, adding cooks to the maximum wins 15/18 market×seat cells (a bottleneck cook unlocks ~$1000–$2800 of contribution for a $60–$180 wage; 10–40× marginal return); max staffing is the profit-max cook count in all three markets at seats ≥30. The overstaffing penalty is real but only appears at seats ≤25. Gate ruling: does not block — understaffing is costly, overstaffing eventually loses money, staffing is not removable as a decision — but staffing level does not vary by market, which reinforces the Domain-B finding that context matters less than claimed. Recommend the playtest script include a small-capacity (~20–25 seat) scenario. H. Regression & scope (Domain D — PASS)Independently reproduced: 117 tests (Core 54, Determinism 31, Scenario 32), 0 fail/skip/warn, ×2 identical; Debug vs Release checksums byte-identical; accounting reconciles to the cent (contribution = revenue − ingredients − labor − overhead, verified incl. an 89-failure service); H1 attribution preserved (no 88%-util flip); integer-cents parsing, FIFO, labor, float/wall-clock/RNG guards all intact; scope clean (no M1/persistent-world code, no committed binaries, bin/obj/assets gitignored). Goldens were re-baselined for a genuine intentional change (D-020/D-022), not to suppress a behavior change. Docs are otherwise honest (cross-OS determinism correctly scoped as OPEN). One low nit: ADR-003 phrases portable determinism a touch more strongly than ADR-002. I. Human-test rulingNOT authorized. Two defects would contaminate the human M0 verdict: the player-facing forecast recommends the wrong pricing direction in a core market (teaching false cause-and-effect), and the balance claim the gate would be entered under ("no dominator / read-the-market-and-flip-regimes") is false (a premium generalist is broadly near-optimal and wins the value lunch). Neither is a determinism/accounting/scope defect — the technical foundation is sound — but both must be corrected or honestly re-scoped, then re-reviewed, before uncoached testers are exposed to the loop. J. Final gateRequired corrections before the human gate (re-review after):
K. Required owner decisionHoward & Aaron must choose the path before the human gate: (a) direct the builder to fix — align the forecast's price response with the sim, and change the economics so premium does not broadly dominate the value lunch — then re-review; or (b) accept an honestly re-scoped balance claim (premium generalist is broadly strong; context tunes not reverses) plus a disclosed forecast pricing-caveat, and re-review a corrected-claims PR before the gate. Either way the two Highs must be resolved and re-verified before the five uncoached M0 playtests begin. Do not merge PR #1 into main until that gate passes; M1 stays unauthorized. — Independent Reviewer (read-only; no branches modified, no PR merged, no human tests run, M1 not begun) |
Second bounded correction inside M0. Not M0.5, not M1. Closes the two findings the first pass left open. Do not merge — for independent re-review, then owner-run human playtests.
H2 / NEW-1 — strategy integrity (price elasticity)
PriceModel: willingness-to-pay anchored to each dish's own suggested price (not the menu median), widened by segment tolerance + dish quality; smooth hyperbolic resistance above WTP, segment-scaled, floored at 3%. No hard cap, no name/recipe branch, no cliff.PriceResistScaleBp=450from an elasticity target, not the dominance outcome.H3 — forecast direction
Verification
docs/design/PRICING-CONTRACT.md,docs/design/FORECAST-CONTRACT.md. Full report:reports/m0/M0-CORRECTION-2-REPORT.md.Disclosed M0.5 residuals (not hidden)
Recommendation: technical Conditional-Pass pending the owner-run human playtests. No M1 work; the Builder does not declare M1 readiness.
🤖 Generated with Claude Code