M0 menu-responsive demand rewrite (Conditional-Pass; owner decision on residual) - #7
M0 menu-responsive demand rewrite (Conditional-Pass; owner decision on residual)#7HSpector1 wants to merge 6 commits into
Conversation
… prototype Design lock for the menu-responsive demand rewrite (per the authorization). Records the failed prior fixed-arrival assumption, the new consideration->captured-demand model sequence, scope/non-goals, and the strategy-integrity + positioning + forecast gates. Honestly records that a v1 prototype (price-level fit x identity fit + mix-weighted draw) FAILED on held-out seeds (mixed dominator still won all markets; focused plans got weaker) and was reverted. The load-bearing risk is the coupled occasion/positioning-coherence deterrent + calibration, which exceeds a single session. Branch + contract + evidence are the handoff; no rewrite PR opened (not ready). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ume + composition) Replace menu-independent arrival composition (the root of the cross-market dominator) with DemandModel.Capture: per-segment consideration from relevant-option depth, positioning focus, occasion variety (BreadthAffinityBp), and a representative price-position gate. Consideration sets BOTH arrival volume and composition; MakeParty now samples the captured mix; the forecast reuses the same Capture. Purchase (PickBest) is unchanged. Quality-scaled depth saturation makes discerning segments need more depth (a lone anchor is not an enthusiast destination). Lunch pool aligned to its value-heavy identity (65→82% value, 120→150 arrivals). Focused Value rebuilt as a full-throughput value operation. Weak-demand attraction floor recalibrated (70→35% of pool) for the new demand regime. Integer/deterministic; no floats/Math.random in Core. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add --searchbase so the dominance search can run on fresh held-out seed bases (the tuning base gets exposed by iteration). Add --probe/--manip/--attr/--fdir diagnostics used to design and validate the demand model; these are the exact experiments an independent reviewer repeats (see M0-DEMAND-REWRITE-REPORT §H). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
New DemandModelTests locks: consideration range/determinism/smoothness, composition responsiveness, the §15 manipulation suite (filler bounded, never-ordered zero-influence, lone anchor partial, broad != full capture), the per-market regime gates (value wins lunch, premium wins enthusiast, mixed viable in social), and the HONEST KNOWN-RESIDUAL (a premium-anchored generalist stays a competitive all-rounder; the value lunch is the pole that resists it). Recalibrate AttributionTests seat/cook regimes and the audited forecast-direction test (assert agreement, not a fixed direction) for the new demand. Re-baseline the three golden checksums (D-030). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Final report (reports/m0/M0-DEMAND-REWRITE-REPORT.md, §A-H) and iteration log (reports/m0/demand-rewrite/candidate-1.md) with the held-out robustness sweep and the decisive economic-root-cause argument. Contract implementation-status updated (Conditional-Pass). DECISION-LOG D-029 (distinct regimes achieved; universal-generalist gate borderline, economic cause; owner decision) and D-030 (goldens re-baselined). Regenerated balance/determinism/forecast evidence (dominance-search on a held-out base). Gate: Verdict Conditional-Pass / Action Continue. Do not merge; no human tests; no M1. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Independent verification (VERIFIED WITH CAVEATS) flagged that dominance-search.md hardcoded "held-out base 900000" even when --searchbase overrode it, and that --seeds does not affect the (fixed-80) dominance search. Print the actual searchBaseSeed and note that --seeds is search-independent; add a NOTE that the best-generalist regret is seed-base-dependent and straddles 10% (read the distribution, not one run's binary verdict). Fold the pooled 7-base regret sweep + the independent-verification result into the report §C. No model/behavior change. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Final Independent Adjudication — PR #7 (M0 Menu-Responsive Demand Rewrite)A. Review identity
B. Executive verdict
C. Rewrite correctness (§5) — CONFIRMED (Sub-Agent A + lead)The funnel is genuinely staged in code: pool = D. Manipulation results (§6) — PASS-WITH-NOTES (Sub-Agent B + lead --manip)All required cases bounded and economically reasonable: one/several cheap fillers keep value consideration at floor (100→343→395 vs pure-value 7399); a never-ordered dish is byte-identical to no dish (7399/8845/3413); a lone premium anchor gives partial (E 3413→5483 < pure-premium 7784) and costs value diners; broad menu under-draws every specialist; two-item menus degrade smoothly; extreme price spread reads as incoherent (all segments depressed); same-average/different-distribution is NOT treated identically (depth+variety respond per-dish) — refuting a plausible concern. Notes: (N1) a cheap filler on a premium menu can inflate social consideration slightly above a genuine social menu (BreadthAffinity quirk; self-defeating at purchase) — Low; (M3) duplicate E. Search evidence (§7,8,9,10,11,12)Methods: shipped harness dominance search across fresh bases (Sub-Agent C) + an independent equal-compute multi-start hill-climb I wrote (pool ~6140, 6 restarts × 160 steps, EQUAL budget for frontiers and generalist, held-out bases), plus a robustness perturbation sweep. Worst-market regret is seed-dependent and straddles 10% — pooled across builder, verifier, Sub-Agent C, and my equal-compute runs: ~5%, 5%, 6%, 6%, 8%, 8%, 9%, 13%, 14%, 17%, 17%, 20%. Dominator (≤10%) on roughly half of held-out bases; no hard dominator anywhere (never <2%, the old failure) and no robust clean pass. Worst market is enthusiast-evening in every case. §8 unequal compute — CONFIRMED (High): the shipped harness hill-climbs each frontier 80 steps but the generalist 120 steps ( §11 robustness — DECISIVE (lead): even where the equal-compute generalist reaches 5% everywhere, it is a fragile knife-edge: only 1/12 (base 525252521) and 0/12 (base 636363631) two-mutation perturbations stay within 10% in all three markets. It is not a broad safe basin — a player cannot casually default into it and small changes cost >10% in some market. §12 structural distinctness — REFINED (Medium, C3): the sharp split is lunch (value) vs dinner (premium) — the lunch frontier differs on ≥4 load-bearing dimensions (menu, price regime, expected meal cost, capacity). But the social and enthusiast frontiers largely overlap (both premium-anchored Roast Chicken/Ribeye/Risotto, differing mainly in price level and seats). So the honest picture is two sharp regimes plus a softer social↔enthusiast gradient, not three fully-distinct regimes. F. Economic decomposition (§13,14) — rebalance JUSTIFIED (Sub-Agent D + lead, independently agree)
The "~5× a value cover" claim is confirmed (revenue and contribution), and per station-minute (the kitchen — the forecast's own named binding ceiling) premium leverage is only G. Player-strategy judgment (§15,16) — healthy generalist, but currently illegibleAgainst the §15 checklist the generalist reads much closer to a healthy generalist than a strategy-collapsing soft dominator: it is never the best (wins no market under strong search), specialists beat it by a material, seed-dependent 5–20%, it carries real operational commitment/risk (a full premium kitchen: high wages, slow dishes), and its advantage disappears under perturbation (0–1/12 robust). The one caveat cutting the other way is legibility (H2): because the positioning/attraction signal is invisible (§H), a player using rough comparisons could default to the premium build without understanding the tradeoff — which would make an otherwise-healthy generalist function as a soft default in practice. Fixing legibility is therefore the highest-leverage change for the product hypothesis. H. Forecast and explanation (§17,18) — CONCERNS (Sub-Agent E + lead, both High reproduced)Forecast directionality is sound (existing
Net: the product question (§2) asks for understandable strategies with visible consequences. The consequences of positioning are currently invisible, so a human playtest would test a game where positioning silently matters but cannot be understood — not an honest test of this design's central claim. I. Regression and scope (§19) — CLEANAccounting reconciles; determinism holds (Determinism suite 31/31; identical checksums across runs); goldens re-baselined and documented (D-030); RNG streams isolated; forecast immutable; causal attribution logic intact (thresholds recalibrated, cause strings not hashed → checksum unaffected); exact-money parsing, FIFO, labor, numeric guards untouched; assets remain quarantined; changed surface confined to demand seams + fixtures ( Key findings (severity format)J. Final gate
K. Required owner decisionHoward & Aaron: authorize the builder to make the two scoped corrections — (a) make per-segment attraction visible in the forecast/autopsy, and (b) course-scale seat/eat duration so multi-course premium meals occupy seats longer — and then return for a re-review that decides the human-test gate? The alternative is to run the 5-player human gate now with those two as documented limitations. Recommendation: do the two corrections first — without (a) the playtest cannot evaluate the very mechanism this milestone adds, and (b) directly attacks the economic driver of the residual soft-dominator without any gate-chasing. Neither correction is M1; both are in-scope M0 finishing work. Read-only review complete. No branch modified or merged. No human tests run. M1 not begun. |
What this is
The authorized headless demand-model rewrite following the PR #5 adjudication. Replaces menu-independent arrival composition (the root of the cross-market dominator) with a menu-responsive consideration model: the menu/prices shape who visits and how many, while purchase (WTP/affordability in
PickBest) stays separate and unchanged.Do not merge. This is a review + owner-decision surface. No human tests. No M1.
What it achieves
DemandModel.Capturederives per-segment consideration from relevant-option depth + positioning focus + occasion variety (BreadthAffinityBp) + a representative price-position gate, and uses it to set both arrival volume and composition (MakePartysamples the captured mix; the forecast reuses the sameCapture).Result — the per-market frontiers are now distinct regimes, which the old fixed-composition model could not produce:
Manipulation-resistant (cheap filler bounded, never-ordered dishes zero-influence, a lone anchor gives some not full premium draw, a broad menu does not capture every segment at full strength — all price-level, never a variance/breadth penalty), legible (plain-language causes), deterministic. 139 tests pass (54 Core + 31 Determinism + 54 Scenario, incl. 14 new demand tests).
The honest residual (owner decision needed)
The strict "no plan within 10% of the frontier in every market" gate is borderline-missed on held-out seeds: a premium-anchored generalist's worst-market regret is ~6–13% (median ~9%, sweep over 4 seed bases; dominator on 2/3 fresh bases).
Decisive cause is economic, not a demand defect: consideration is market-name-agnostic, so a premium menu serves the same fraction of whatever premium slice a market's pool holds, and premium multi-course checks (~5× a value cover) make that profitable everywhere. The regret is fixed by (pool mix) × (per-cover economics) — both outside the demand model — so no composition candidate (C1–C4) can close it. Closing it robustly needs an economic rebalance of value-vs-premium margins, or the excluded systems (reputation / repeat visits), or an extreme value-heavy lunch pool. One principled fixture change was made (lunch pool 65→82% value, matching its stated "value-heavy" identity, achieving 3 of 4 gates); pushing to ~95% to force the 4th was declined as gate-chasing.
Gate recommendation
Owner decision: (a) accept the borderline all-rounder (specialists still win each pole — arguably healthy), (b) authorize a targeted economic rebalance follow-on, or (c) accept a strongly value-heavy lunch pool.
Where to look
reports/m0/M0-DEMAND-REWRITE-REPORT.md— full report (§A–H incl. exact experiments to re-run)reports/m0/demand-rewrite/candidate-1.md— iteration log (R0–R6, failed square-sharpen preserved)docs/DECISION-LOG.md— D-029 / D-030--probe,--manip,--attr,--fdir,--searchbase🤖 Generated with Claude Code