M0: expose segment attraction and scale dining time by courses - #9
M0: expose segment attraction and scale dining time by courses#9HSpector1 wants to merge 5 commits into
Conversation
Two authorized corrections only: H-LEG-1 (attraction legibility — expose per-segment consideration/captured mix in forecast+autopsy, add weak-attraction/wrong-mix causal categories) and MED-ECON-1 (course-scaled seat duration — shared SeatEatMin primitive, Model A base+per-distinct-course, sim + forecast). Locked before coding per §6. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
MED-ECON-1: replace flat EatMin=12 with a shared Tuning.SeatEatMin(coursesX10) (base 6 + 6 per distinct course, clamped [12,30]); the sim scales a party's eat window by its DistinctCourseCount and the forecast scales seat-cover capacity by the captured mix's expected course count, so a premium multi-course meal turns fewer tables (seat-minute leverage ~4x -> ~3x). Only eat time changes; ticket/cook path untouched. H-LEG-1 (core): thread CapturedDemand into Finalize; build realized per-segment attraction (pool/consideration/arrival-mix); add weak-attraction, wrong-mix, and purchase-conversion causal categories with disjoint-PState attribution (wrong-mix is a composition opportunity cost, never double-counted). (Note: Simulator.cs/Forecast.cs carry both corrections; presentation is in the next commit.) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…tation) Add SegmentAttraction (pool/consideration/composition/parties) to ForecastSnapshot and ServiceResult; render a "Who this plan attracts" table pre-service with a plain-language attraction note (largest-remainder reconciliation), and a per-segment ATTRACTION funnel post-service, so the menu-responsive demand model is visible and forecast-vs-actual mix can be compared. Harness --report diagnostic to eyeball forecast+autopsy. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…and re-baseline goldens New ReadinessTests: SeatEatMin monotonic/bounded, ticket-time unchanged, per-segment attraction reconciles pre/post-service, wrong-mix reachable (and never for a good premium plan), weak-attraction and seating remain distinguishable. Rename the weak-demand attribution assertion to weak-attraction. Re-baseline the three goldens (course-scaled duration changes seat occupancy; attraction legibility does not affect the checksum). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Final report (reports/m0/M0-HUMAN-GATE-READINESS-REPORT.md, §A-J), DECISION-LOG D-031 (the two corrections) and D-032 (goldens re-baselined). Regenerated balance/determinism/ forecast evidence; the example autopsy now renders the pre-service attraction table and the post-service ATTRACTION funnel. Gate: Verdict Pass-with-notes / Action Continue (request focused independent re-review). Do not merge; no human tests; no M1. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Final Focused Independent Re-Review — M0 Human-Gate Readiness (PR #9)
A. Review identity
Baseline confirmed: the diff ( B. Executive verdictBoth authorized corrections are implemented, in scope, deterministic, and independently verified. No Blocker or High defect survives scrutiny. The single Blocker raised by a sub-agent (attraction table "attracted > pool") was refuted on cross-check and re-adjudicated by the lead to a Medium legibility note: it is internally consistent with the simulator, matches the authorizing contract's exact formula, and does not teach a false strategic rule. The remaining items are legibility/wording edge cases and one economic-rationale disclosure — none makes attraction invisible or corrupts the funnel the gate relies on. This maps to §31 "Authorized with stated limitations." Load-bearing claims reproduced independently (held-out seeds):
C. Attraction legibility (H-LEG-1)Forecast table ( Actual funnel (autopsy): a Player questions (§26) — answerable from the CLI alone:
Terminology issues: "appeal" for consideration and "mix" for composition read cleanly. The one confusing point is the D. Causal attributionThe autopsy's causal summary distinguishes five disjoint channels and reports the single largest by measured lost contribution (ranked by
Disjoint-state audit (§11): every generated party ends in exactly one terminal PState (Paid / WalkedSeat / WalkedFood / NoOrder); Controlled fixtures (lead, 200 held-out seeds each):
wrong-mix is reachable but never fires for a well-matched premium plan (0/200), and only 0.5% for a healthy value plan (not a systematic false positive). Because M0 is throughput-dominated (D-012), "kitchen" is the primary cause in most services; the demand-side causes are narrow-regime by design, but the attraction table is always shown, so legibility does not depend on which cause fires. Priority (§15): strong-attraction + weak-op → operational primary; weak-attraction + excellent-op → attraction primary with no add-capacity advice; weak + weak → both stages reported and ordered. Verified. E. Timing correction (MED-ECON-1)
F. Economic effects & premium-leverage rulingPer-seat/per-station decomposition (base 555000000, n=200), premium retains every qualitative property: check $58.59 vs value $20.07; ingredient/cover $14.92 vs $4.53; ticket 24.2 min vs 11.8; effective dwell 65 vs 33 min; station util 63-64% vs 43-46%; contribution/cover $34.28 vs $9.94. Premium still wins the enthusiast evening; value still wins lunch; mixed viable in social. No over-correction; no correction that is economically inert. Covers/seat clarification (§22): the builder's 3.67 (value) / 2.96 (premium) are theoretical seat-turn figures (= ServiceMinutes ÷ dwell) measured in controlled 24-seat seat-bound rooms — a basis the builder's report explicitly discloses (reproduced independently: 3.69 / 2.77). They are not realized covers/seat, which for the named strategies in their native markets are ≈ 3.21 / 2.51. The gate summary should carry this label so the figure is not over-generalized. Premium-leverage ruling (§23): Credible with tuning notes. G. Strategy regression
H. Technical regression & scope
I. Findings(Severity: Blocker/High/Medium/Low/Note · Status = whether the tested claim held) A1 — Forecast "attracted" can exceed a segment's "pool"
A2 — Over-attraction note is silent in the unstaffed-station case
A3 — Self-contradictory note for a menu with no main course
A4 — Note fixates on a micro-segment for premium plans · Low · Confirmed
A5 — "largest-remainder" label is really "remainder-to-last" · Note · Confirmed
B-2 — Loss-making seating verdict prints "~$0.00 of lost contribution" yet advises adding seats · Low · Confirmed
B-3 — purchase-conversion is a thin residual channel · Note · Refined
B-4 — wrong-mix cents is a modeled opportunity estimate · Note · Confirmed
D-3 — Forecast ExpectedCovers can point the wrong direction vs the sim on seat/course changes · Low · Refuted (as a timing defect)
D-4 — Seating-risk note is a kitchen-overload warning, not a seat-shortage warning · Note · Refined
E-1 — Premium-leverage rationale is metric-dependent and partly inverted · Medium · Refined
E-2 — 3.67/2.96 covers/seat are controlled seat-turns, not realized · Low · Refined
F-DOC-1 — Stale code comment: "~70% / 62%-converting" vs the actual 35% floor · Low · Confirmed
Confirmations that support readiness (Note/Confirmed): B-1 (conservation, 56,700+ runs, 0 violations), B-5 (one-course byte-identity + determinism), C-1..C-5 (timing primitive, distinct-course logic, seat coupling, table release, exact-2× dwell), D-1/D-2/D-5 (shared primitive, seats-past-kitchen-wall flat, ticket-time isolation), E-3/E-4 (generalist regret >10%, determinism/test integrity), F-SCOPE-1 (clean scope), F-TEST-1 (no test weakening). J. Human-test rulingAuthorized with stated limitations. Both corrections work; the remaining imperfections (A1 label, A2/A3 note edge cases, A4 note priority, E-1 rationale) do not teach false strategic rules; the generalist is a healthy fragile strategy; players can understand and test positioning from the CLI; only minor tuning/wording notes remain. Recommended (not required) before the gate: a small notes-polish pass (A1 caption, A2 unstaffed-kitchen warning, A3 no-mains phrasing) would raise fidelity on the very feature the gate tests — owner's call, since the core table is legible and the degenerate cases already surface loud 0-cover / negative-contribution signals. K. Final gate
Keep PR #9, the rewrite/correction chain, and the foundation unmerged until the human gate result is in; M1 remains separately unauthorized. Independent review complete. Read-only throughout: no branch modified, no commit, no merge, no human tests run. This review reproduced every load-bearing claim on held-out seeds; where it disagreed with a sub-agent (A1 severity, E-1 status) the lead re-adjudicated with cited evidence. |
What this is
The two authorized M0 finishing corrections from the PR #7 adjudication (H2 legibility + M1 economics). Focused, in-scope; does not reopen
DemandModel.Capture. Targets the PR #7 branch for a clean diff.Do not merge. Review surface + human-gate decision. No human tests. No M1.
H-LEG-1 — attraction legibility
The menu-responsive demand model is now visible.
SegmentAttraction(pool / consideration% / mix% / parties) is onForecastSnapshotandServiceResult:MED-ECON-1 — course-scaled seat duration
Flat
EatMin=12→ sharedTuning.SeatEatMin(coursesX10) = clamp(6 + 6·courses, 12, 30)→ 1/2/3-course = 12 / 18 / 24. The sim scales a party's eat window byDistinctCourseCount; the forecast scalesseatCoverCapby the captured-mix expected courses (same primitive). Only eat time changes —AvgTicketTimeMin(kitchen) is untouched.Economic effect (empirical): a premium 3-course service turns 2.96 covers/seat (~61-min dwell) vs a value 1-course 3.67 covers/seat (~33-min dwell) — premium ties up a table ~1.85× longer per seat-hour, so per-seat-minute leverage falls from ~4× toward ~2.5–3×. No over-correction: premium still wins enthusiast, value still wins lunch, mixed viable in social, KNOWN_RESIDUAL holds. Favorable side effect: best-generalist regret rose 6%→17% on base 271828182 (exactly the mechanism the adjudication predicted), not gate-chasing.
Technical
Gate recommendation
Continuedoes not authorize human tests, M1, or merging.Where to look
reports/m0/M0-HUMAN-GATE-READINESS-REPORT.md— §A–J incl. exact reviewer experiments (§J)docs/design/M0-HUMAN-GATE-READINESS-CONTRACT.md— the pre-code design lockdocs/DECISION-LOG.md— D-031 / D-032--report(attraction table + causal categories),dotnet test🤖 Generated with Claude Code