Foundation + M0 Headless Service Lab (stops at M0 gate) - #1
Conversation
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…docs Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…sset decisions Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…autopsy, revise) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ty tests (74) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…tion Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Independent M0 Audit — PR #1A. Review identity
B. Executive verdict
C. Claims table
D. Blocking findingsNone. Nothing makes the M0 lab untrustworthy as a deterministic, in-scope, reconciling simulation; nothing is fabricated; determinism holds; the build and tests are honest. E. Non-blocking findingsHighH1 — Causal autopsy misattributes the bottleneck in a boundary band (AUTOPSY-1). H2 — "No dominant strategy" is overstated; a dominant configuration exists (BAL-1). H3 — Forecast is systematically biased high; its confidence band almost never contains the outcome (F1). Medium
Low (12) and Note (20)Selected: DET-3 checksum covers economic aggregates only (excludes forecast + diagnostic fields) — state the scope so M1 doesn't read "golden stable" as "whole result unchanged". DET-5 F. Independent experiments (actual results)
G. Architecture assessment
H. Simulation assessmentMenu/pricing/staffing/capacity produce genuinely different, context-dependent outcomes; throughput, staffing, pricing, capacity, failure, and revise-and-rerun all work and reconcile. Quality/causality are the weak spots: quality has no economic teeth beyond comped failures (confirmed), which enables the dominant archetype (H2); and causal attribution is correct in the clear cases but wrong in the boundary band (H1). Recovery/revision works mechanically; the loop's fairness depends on fixing the forecast (H3) and autopsy (H1). I. Balance assessmentReproduced: 2/3 distinct named winners; deliberately-bad strategies all lose and are (mostly) legibly diagnosed; no fixture leakage — the sim contains no branch on J. Scope assessmentEvery hard non-goal: Absent in executable code (graphics, engine, pathfinding, spatial movement, inventory, supplier, spoilage, persistent save, database, campaign, competitor, city, manager, delegation, franchise, multi-restaurant, marketing, critic, audio, animation, Steam, telemetry, mods, asset imports). Deferred systems appear in docs only. Asset packages quarantined and not committed (no binaries tracked). Clean. K. Test-quality assessmentThe suite proves: ledger reconciliation, funnel non-negativity/ranges, comp-removes-failed-revenue, same-seed determinism, no-instance-float, no-wall-clock/ L. Human-test readinessThe package can mechanically support a playtest: the CLI exposes inspect → plan → forecast → commit → autopsy → revise → run again, and the script/sourcing/consent docs are complete. But it is not yet fair to run: testers would be misinformed by the miscalibrated forecast (H3) on every run and misdirected by the boundary-band autopsy (H1). Fix or explicitly disclose both before the five-player gate, or the human evidence will be contaminated. M. Required bounded correction pass (corrections only, no new features)
N. Final gate recommendationConditional-Pass (not Pass-with-notes) because H1 and H3 are real behavior/correctness issues — not cosmetic — that must be corrected before the autopsy and forecast are trusted. Defer because the five-player human gate is outstanding and M1 is not authorized; do not read a technical pass as authorization to continue into M1. O. Required owner decisionHoward & Aaron: authorize a single bounded M0 correction pass (fix the autopsy misattribution H1 + forecast calibration/disclosure H3 + correct the "no dominant strategy" and cross-OS claims) before running the five uncoached human playtests — versus proceeding to playtests now with those caveats explicitly disclosed to testers. Either path is legitimate; choosing "proceed with caveats" means accepting that the forecast and boundary-band autopsy may bias what testers report. Merge recommendation: do not merge as-is; merge after the section-M corrections (they are corrections, not features). M1 remains unauthorized independent of the merge decision. The PR branch was not modified during this review. |
M0 Headless Service Lab — foundation
Establishes the repository as the project's source of truth, folds in the accepted review findings, and builds the M0 Headless Service Lab. Stops at the M0 gate. Do not merge as "M1 authorized" — this closes only the technical half of M0; the fun/comprehension/replay half needs owner-run human playtests (see below).
What's here
docs/archive/2026-07-28/; an activeMASTER-PLAN.mdchange-record; commercial hypothesis + go-to-market + comparables + Early-Access decision; honestEFFORT-AND-SCOPE(through-M3 = first-commercial cut; full ceiling is multi-year); a risk register with human/commercial risks (burnout, founder deadlock, funding, motivation, commercial indifference); determinism contract + 4 ADRs; fantasy-continuity resolution + deferred proof gate; playtest sourcing/script/consent; asset provenance + quarantine.src/RestaurantSim.Core): deterministic, integer-only, seeded RNG streams, no wall clock. Demand→choice→seating→kitchen→FOH→satisfaction→economy→causal autopsy→forecast→checksum. 3 segments, 12 recipes, 8 employees, 3 markets, 9 named strategies. CLI decision loop + distribution/determinism harness.Evidence (in
reports/, 200 seeds/cell)Scope audit
Clean. Nothing from the M0 non-goals list was built; asset packages quarantined and not committed.
Gate recommendation (
reports/m0/M0-GATE-RECOMMENDATION.md)Conditional-Pass (technical M0 proven) · Defer the Continue/Rewrite/Abandon decision to the owners pending an independent review + five uncoached human playtests. M1 is not begun and not authorized.
🤖 Generated with Claude Code