SPEC_VERSION: 0.3.1. The decisions behind it are in docs/adr/.
This document is the single source of truth for every cross-workstream contract.
If code and SPEC disagree, the code is wrong. Changes require a spec-change
issue, a PR titled spec: ... (or spec!: ... if breaking), and approval from
the human owner (see CODEOWNERS). Breaking changes bump the minor version; a
change that moves no contract bumps the patch version. Either must update
SPEC_VERSION here and the spec_version constants in code.
-
ISA: RV32IM, little-endian. No FPU (DOOM is fixed-point). No compressed (C) extension — compile with
-march=rv32im -mabi=ilp32. -
Privilege: machine-mode only, flat physical addressing, no MMU, no interrupts. All device interaction is polled MMIO (§3).
-
ecall: fatal halt, reason =ECALL.ebreak: fatal halt, reason =EBREAK. Any CSR instruction (csrrw/csrrs/csrrc/csrrwi/csrrsi/csrrci): fatal halt, reason =CSR. None should ever appear in the ROM; if they do, the ROM is misbuilt. -
fence/fence.i: not a fatal halt — a plain retiring no-op (decodes onto the same collapsed arm as a real no-op, e.g.addi x0,x0,0;pc += 4, no other state change). A single hart with no cache has nothing to reorder against, and the toolchain emits these from compiled C — halting on them would break the ROM. Agreed cross-engine in issue #37; no automated test currently exercises this (riscv-tests'fence_i.Sis excluded, since it also exercises self-modifying code, which ADR-0002 forbids outright — see #37), so treat "does my decode give FENCE its own arm, or does it fall through toILLEGAL_INSN?" as a thing to verify by reading the decode path, not by memory. -
Misaligned word/halfword data access (load/store): fatal halt, reason =
MISALIGNED, with pc and target address in the halt record (compile ROM with-mstrict-align, keeping the SQL load/store path branch-free). -
Misalignment is decided before the address's region and before the access's width. An access that is both misaligned and outside every region of §2 halts
MISALIGNED, notBAD_ADDR; so does a misaligned narrow store to FRAMEBUFFER or PALETTE, which §2 clause 2 would otherwise makeBAD_ADDR. This matches the RISC-V privileged architecture's synchronous-exception priority, which ranks address-misaligned above access-fault. Both rules can fire for one access andhalt_reasonis observable state both engines must produce identically (§1's closed vocabulary), so the order is part of the contract rather than a property of each engine's arm order. -
Misaligned jump/branch target (
jal/jalr/a taken branch computes a target not aligned to 4 bytes): fatal halt, reason =MISALIGNED, checked at the transferring instruction, eagerly — matching real RISC-V's instruction-address-misaligned semantics, where the exception is reported on the branch/jump, not on the target. The halt record's pc is the transferring instruction's own pc (not the unreached target), and neitherpcnorrdis updated — the jump does not architecturally complete. Ruled in issue #37 over the alternative (deferring the check to the next fetch, with the target as the halt's pc): the eager check keeps the halt inside the instruction that caused it, and removes an ordering dependency between engines over exactly when "the next fetch" happens relative to batch boundaries. In practice unreachable from a correctly-built ROM (no compressed-instruction extension means every real target is 4-byte aligned, andjalrclears bit 0) — which is exactly why it needs pinning rather than being left to each engine's judgment: an unreachable path never gets exercised by a passing test, so the written agreement is the only thing keeping two independent engines aligned on it. -
Unimplemented/illegal opcode: fatal halt, reason =
ILLEGAL_INSN, with pc and raw instruction word in the halt record. -
Instruction fetch from a pc outside the text region
[text_start, text_end)(§2, §4): fatal halt, reason =BAD_ADDR, with the pc in the halt record. The text region is the only part of RAM an engine is required to be able to fetch from: the executor pre-decodes it into a table (ADR-0002) and holds nothing for the rest of RAM. Unreachable from a correctly-built ROM, which does not transfer control outside its own text, and pinned for the same reason §1's eager misaligned-target check is: an unreachable path never gets exercised by a passing test, so the written agreement is the only thing keeping two independent engines aligned on it. The check is on the fetch, so the halt record's pc is the unfetched address itself and no instruction retires. -
Store into the text region (§2): fatal halt, reason =
SELF_MODIFY, with pc and target address in the halt record. The executor pre-decodes text into a table (ADR-0002) because decoding inside the fold costs 7.4× the throughput; a write to text would silently invalidate that table. DOOM does not self-modify — if it appears to, that is a bug worth halting on. -
Halt-reason vocabulary is closed and exact-match:
ILLEGAL_INSN,BAD_ADDR,SELF_MODIFY,MISALIGNED,ECALL,EBREAK,CSR,EXIT— all uppercase ASCII, no punctuation, no per-engine normalization. It is observable state (cpu_state.halt_reason, §5) that bothrefemuandsqlcpumust produce identically; agreed cross-engine in issue #37. The first seven are faults;EXITis not. It is the ROM's own clean stop, written to §3'sEXITregister, and it appears here because §5'scpu_stateroutes every way the machine can stop through this one column. "Halted normally" is thereforehalt_reason = 'EXIT', with the written value inexit_code— not an empty string, which would be indistinguishable from an unset column in a differential comparison. A fault never setsexit_code, andEXITnever sets a fault reason, so the two are always separable. This matters at exactly one moment and it is an expensive one:EXITis how-timedemo demo3terminates, so it is the last value a victory run produces, and a cross-engine mismatch here fails the §7 final-state comparison after everything else has already agreed. No riscv-test writes that register.A conforming
-timedemo demo3completion'sexit_codeis4,294,967,295((uint32_t)-1) — not a failure signature despite the value's usual connotation.-timedemo's completion path is notI_Quit; it isG_CheckDemoStatus'sif (timingdemo)branch callingI_Errordirectly, which falls throughZenityErrorBox→ZenityAvailable→system(ZENITY_BINARY " --help ...")— a syscallrom/src/syscalls.cnever implements — before its ownexit(-1). Pinned from a measured, complete run, not the isolated mechanism probe that first found this path (issue #107/#111):refemuran the fulldemo3timedemo to termination (issue #129, againstPINNED_HASH eabb12ed…) — 2,836,207,097 instructions, halting cleanly withhalt_reason = 'EXIT',pc = 0x8000_06b0(the same_exitMMIO-write-then-spin address the isolated probe found),exit_code = 4,294,967,295— confirming the probe's mechanism holds for a real, full-length run and not only in isolation. Re-measured after #175's instruction-count optimization changed which specific binary is pinned (issue #202, againstPINNED_HASH 9a6a47d0…): 2,300,210,133 instructions — a different length, consistent with #175's own accounting, not a discrepancy — but the samehalt_reason,pc, andexit_codeas the first run, confirming the mechanism isn't an artifact of one specific binary. Worth pinning explicitly rather than leaving as "whateverEXIThappens to carry": the exit path depends onsystem()continuing to resolve the way this newlib build currently does (returning promptly rather than hanging or erroring differently), which a future change torom/src/syscalls.ccould alter without anyone intending to change what "victory" means (issue #121). A-timedemo demo3run's success is therefore judged onhalt_reason = 'EXIT'andexit_code = 4,294,967,295together, alongside §7'sfbhash— ahalt_reasonmatch with a differentexit_codeis a real divergence, not a benign difference. -
Reset state:
pc = 0x8000_0000, allx1..x31 = 0.crt0setssp, zeroes.bss, and jumps tomain.x0hardwired to 0 (obviously — but the SQL register file must enforce writes to x0 being discarded). -
A fatal halt does not retire the instruction that caused it:
icountis not incremented, and no architectural state (pc,rd, memory) is modified. The halt record'spcidentifies the faulting instruction.icountis load-bearing here, not cosmetic — §7's checkpoint trace and §3.1's elastic time both key on it. This applies to every fatal halt regardless of which section defines it — includingBAD_ADDR(§2), not only §1's list above. (#72 — ruled after sqlcpu's riscv-tests harness and refemu/executor disagreed onicountby exactly one on every fixture.)
| Region | Base | Size | Notes |
|---|---|---|---|
| RAM | 0x8000_0000 |
24 MiB | ROM image loaded at base; code+data+heap+stack |
| MMIO | 0x1000_0000 |
4 KiB | Registers, §3 |
| FRAMEBUFFER | 0x1100_0000 |
64,000 B | 320×200, 8bpp palette-indexed, row-major |
| PALETTE | 0x1101_0000 |
768 B | 256 × RGB (3 bytes), written on palette change |
Anything outside these regions: fatal halt, reason = BAD_ADDR.
Valid access to FRAMEBUFFER/PALETTE is narrower than the table above implies, in two ways that are different kinds of claim and are pinned differently on purpose:
- The ROM shall not read from FRAMEBUFFER or PALETTE. An
implementation MAY treat such a read as a fatal halt, reason =
BAD_ADDR(reusing the reason from the sentence above, not a new one) — but is not required to. This is an obligation on the ROM, not an assertion that the memory itself is unreadable: FRAMEBUFFER/PALETTE are ordinary, physically readable memory that this ROM simply chooses never to read back (§2's own rationale already says why — display output is read back through the SQL-side render query, never a CPU-side load).refemu's referenceMemorykeeps a faithful, well-defined read path for both regions and remains conformant;sqlcpu/executorMAY instead fatally halt on such a read, since a conformant ROM never exercises that path, and a violation is still caught loudly — the executor halts whilerefemureturns data, exactly the divergence differential comparison (§7) exists to surface. Established empirically, not asserted:rom/src/dg_hooks.c— the only code that ever writes to either region — declaresframebuffer_mmio/palette_mmioasstatic volatile uint32_t *const(internal linkage: no other translation unit can even name them) and uses both exclusively as assignment targets; a repo-wide grep ofrom/srcandrom/vendorfor the regions' addresses and pointer names finds no other occurrence anywhere, including the vendored DOOM engine and every doomgeneric platform backend, which touch onlyDG_ScreenBuffer(heap) andcolors[](RAM) and have no concept that either is MMIO at all. An instrumentedrefemurun to icount 60,000,000 (72 committed frames, well past boot and the first frame into real demo playback) observed zero reads to either region (issue #130/#134). - Stores must be word-width: a byte or halfword store to either
region is a fatal halt, reason =
BAD_ADDR, for both engines. Unlike clause 1, this is not about readability and is not optional for either engine — a narrower store is normally a read-modify-write against the word it lands in, and clause 1 means there is no previous value to read: nothing has, or ever will, load from these regions to blend against (whether or not a given engine's own memory model happens to be capable of it). A sub-word store therefore has no correct value to write in either engine; silently dropping it or writing a wrong word are the only alternatives to halting, and neither is honest. This extends an existing pattern rather than inventing one — §3's own header is "MMIO registers (word access only)" — so §2 becomes consistent with §3 instead of carrying a narrower rule than it. It also makes a boundary overrun impossible for free: a word-aligned, word-width store at an in-range address cannot spill past a region whose size (64,000 and 768 bytes) is already a multiple of 4.
Within RAM, the text region [text_start, text_end) is declared by the ROM
manifest (§4) and is read-only: it is the region the executor pre-decodes,
and a store into it is SELF_MODIFY (§1). The linker script places code and
nothing writable there.
Rationale for 8bpp + palette (not doomgeneric's default 32bpp buffer): 4× fewer store instructions per frame on the emulated CPU. Frame conversion to RGB happens in the render query (SQL side — allowed) at readout time.
| Addr | Name | R/W | Semantics |
|---|---|---|---|
0x1000_0000 |
TICKS_MS |
R | Emulated milliseconds = instructions_retired / IPMS (§3.1) |
0x1000_0004 |
KEYQ |
R | Pop next key event; 0 if queue empty. Encoding: §3.2 |
0x1000_0008 |
EXIT |
W | Halt emulation; written value = exit code |
0x1000_000C |
PUTCHAR |
W | Debug console: append low byte to console_out table |
0x1000_0010 |
FRAME_COMMIT |
W | ROM signals framebuffer complete; value = frame number |
Accesses to the MMIO window that are not word-width, or whose address is
not one of the five register offsets above, read as 0 and are silently
ignored on write — no side effect, no fatal halt. Reproducing a
byte-addressable scratch region for this window would cost node-evaluation
budget in the executor's fold on every retired instruction (§6) to serve
behavior no ROM address exercises; DOOM's platform layer only ever declares
these five offsets as volatile uint32_t *. Agreed cross-engine in issue
#87.
IPMS (instructions-per-emulated-millisecond) is a constant in
executor/config, default 10,000 (≈10 MHz virtual CPU). Time advances
with retired instructions, never wall clock. This makes execution fully
deterministic and speed-independent: -timedemo produces identical frames
whether the emulator runs at 1 kIPS or 1 MIPS. Never derive time from
now() or batch wall time.
KEYQ read returns (pressed << 8) | doomkey where pressed ∈ {0,1} and
doomkey is the doomgeneric keycode. Reads pop exactly one event; the queue
is a ClickHouse table (input_queue) populated by the driver, ordered by
event_seq. A read when empty returns 0 and pops nothing.
Deliverable from rom/: two files, built reproducibly (pinned toolchain
image, no timestamps):
doom-rv32im.bin— flat binary, loaded verbatim at0x8000_0000. Contains code, rodata, the embedded sharewaredoom1.wad, and zero-init markers. Note: sharewaredoom1.wadis freely redistributable; the DOOM source is GPL. Do not embed commercial WADs in the repo.manifest.json—{"spec_version", "entry": 2147483648, "load_addr", "size", "sha256", "text_start", "text_end"}. The text bounds are absolute addresses delimiting the read-only region of §2; the build emits them from the linker script rather than having them written by hand.
CI pins the expected sha256 in rom/PINNED_HASH; a mismatch on an
unrelated PR means the build went nondeterministic — treat as P0.
Authoritative DDL lives in sqlcpu/schema.sql; this section defines the
shape. All tables carry spec_version String.
-
cpu_state— one row per committed batch:(batch_id UInt64, icount UInt64, pc UInt32, regs Array(UInt32) /* len 31, x1..x31 */, halted UInt8, halt_reason LowCardinality(String), exit_code UInt32). A durable table, never pruned, derived frombatch_commit(below) by the same idempotent flush that populatesramandconsole_out— a fourth derivation on identical terms, not a special case.ReplacingMergeTreekeyed bybatch_id, so a flush redone after a crash cannot leave two rows for one batch; the derivation is deterministic from a singlebatch_commitrow, so any duplicates would be byte-identical anyway, but "one row per committed batch" should be literally true of the table rather than merely true of what a reader happens to select. Read it withFINALwhen the row count or the full history matters; the per-batch state reload (ORDER BY batch_id DESC LIMIT 1) does not need it, since a duplicate pair is content-identical and either row answers correctly. -
batch_commit— the batch's single atomic write (§6). One row per batch, superset ofcpu_state's columns plus the per-batch bulky data recovery needs to safely re-deriveramandconsole_out:keyq_pos UInt64(cumulative KEYQ pops through this batch),has_frame UInt8/frame_no UInt32(FRAME_COMMIT, if any, this batch),wl_addr Array(UInt32),wl_val Array(UInt32),wl_icount Array(UInt64)(per-store version — see the versioning note below),fb_wl_addr Array(UInt32),fb_wl_val Array(UInt32),fb_wl_icount Array(UInt64),pal_wl_addr Array(UInt32),pal_wl_val Array(UInt32),pal_wl_icount Array(UInt64)(§2's FRAMEBUFFER/PALETTE write-log — same shape aswl_*above, one triple of arrays per region, recovery-needed to re-deriveframebuffer/palettebelow the same waywl_*re-derivesram; #130/#160),console_bytes Array(UInt8), andcp_icount Array(UInt64),cp_pc Array(UInt32),cp_regs Array(UInt32)(§7'sCHECKPOINT_INTERVALcheckpoints, one entry per boundary a retiring instruction of this batch landed on.cp_regsis flat, 31 words (x1..x31) per entry, incp_icount's own order. The checkpoint cadence is finer than a batch andarrayFoldexposes no intermediate accumulator, so a boundary's state is observable only if the fold records it; a boundary landing on the batch's own last retired instruction is the committedcpu_staterow instead. Unlike the write-logs these are not recovery data: they are read once, after the batch, and go with the row when retention drops it). Bounded by retention on batch_id lag, not wall-clock time: only the most recent N rows are kept (N = 16,executor/config), older ones dropped whole by a fixed statement (partition-drop, or a delete keyed onbatch_id < (SELECT max(batch_id) - N FROM batch_commit)) the driver issues unconditionally every batch. Retention drops entire rows rather than individual columns, and touches onlybatch_commit—cpu_stateis derived and kept forever, so §5's "one row per committed batch" survives the window. That asymmetry is the point: the atomic write is short-lived scaffolding, the derived state is permanent — the threshold is computed inside the query, not decided by driver logic, so this stays within PURITY.md's housekeeping allowance. In normal operation only the latest row's bulky columns are ever read (everything older has already been flushed intoram/console_out), so N=16 is generous headroom, not a tight bound. Rejected alternative: a wall-clock TTL (an earlier draft of this PR usedcommit_ts DateTime DEFAULT now()with a 1-day column TTL). Rejected because it reintroduces exactly the failure mode this table exists to avoid: a driver outage longer than the TTL window loses the last committed batch's write-log before recovery can replay it, silently and permanently divergingramfromcpu_statewith nothing to reconcile from. For this project that is not an edge case — ademo3timelapse run is multi-week at realistic throughput, machines sleep, and work routinely pauses on a human gate, so "the outage exceeds one day" is an expected occurrence, not an exotic one. Batch-id-lag retention is strictly better on every axis that mattered: no clock anywhere (thenow()and its purity annotation disappear entirely), no outage hazard (recovery is unconditional, independent of how long the driver was down), deterministic per §8, and a tighter bound in practice than a day's worth of batches. -
ram—ReplacingMergeTree(version)keyed byword_addr UInt32(byte addr >> 2),value UInt32,version UInt64(= icount of the store). Loaded once per batch as a constant array; stores inside a batch live in the fold's write-log and are flushed here on batch commit. The version for each delta must be the individual store's ownicount(batch_commit.wl_icount[i]), not the batch's finalicount— two same-address stores in one batch sharing a version is aReplacingMergeTreetie with an unspecified winner, which violates §8's explicit-ordering rule. -
framebuffer—ReplacingMergeTree(version)keyed byword_addr UInt32,value UInt32,version UInt64(= icount of the store, same convention asram) — persistent storage for §2's FRAMEBUFFER region, mirroringram's shape exactly.word_addrhere is relative to FRAMEBUFFER's own base (0..15,999,(byte_addr - 0x1100_0000) >> 2), not RAM-relative or absolute — there is no RAM_BASE-style rebasing step on this table's flush, unlikeram's (wl_addris RAM_BASE-relative and must addRAM_BASE >> 2back on to becomeram.word_addr's absolute convention;fb_wl_addris already region-relative on both sides of the flush, so no such adjustment applies here). Word-only by construction (§2 clause 2 — every store landing infb_wl_addr/fb_wl_valis already a full-word store, or the executor haltedBAD_ADDRbefore it could retire), sovalueis always the complete word, never a partial blend.DG_DrawFrame's 64,000 bytes divide into exactly 16,000 words with no remainder, so this table'sword_addrrange is dense over[0, 16,000)once every pixel has been drawn at least once — never partially written mid-word the way a byte-decomposition scheme would risk. -
palette— identical shape toframebufferabove, for §2's PALETTE region:word_addrrelative to0x1101_0000, range[0, 192)(768 bytes / 4). Written far less often (on palette change, not per-pixel) — a separate table fromframebufferrather than one combined table with a region tag, so a palette read never pays framebuffer's write volume and vice versa (#130). -
input_queue—(event_seq UInt64, key_event UInt16, consumed UInt8). -
frames_out—(frame_no UInt32, committed_icount UInt64, fb String /* 64,000 bytes */, palette String /* 768 bytes */); written by the render query onFRAME_COMMIT.committed_icountis the icount after theFRAME_COMMIT-writing store itself retires — the same post-increment convention aswl_icount/fb_wl_icount/pal_wl_icountabove (§1: icount counts retired instructions, and the MMIO store that writesFRAME_COMMITis an ordinary retiring instruction, not a fatal halt, so it increments icount like any other). Called out explicitly because it is easy to get backwards by one: a naive read of "the icount whenFRAME_COMMITfired" could mean either the icount before that instruction's own retirement or after it, and only the latter is consistent with every othericountvalue in this schema meaning "count of instructions retired so far, inclusive of the current row's own triggering instruction."#29's target checkpoint is this post-increment value. Itsfb_hash(fe5d82c0f42d45f1) is stable across a conforming binary's exact instruction cost — confirmed by #175's instruction-count optimization, which changed the icount below without changing the hash — but the icount itself is not:15,653,137against the originalPINNED_HASH eabb12ed…,15,393,136against the currentPINNED_HASH 9a6a47d0…(issues #175/#198). -
console_out—(seq UInt64, byte UInt8). -
decoded— the pre-decoded text segment (ADR-0002), built by a SQL query overramat ROM load and covering[text_start, text_end)only:(word_addr UInt32, id UInt8, rd UInt8, rs1 UInt8, rs2 UInt8, imm UInt32, tgt UInt32, mk UInt32, sg UInt8, m_sg1 UInt8, m_sg2 UInt8, m_hi UInt8, d_sg UInt8, cmp_sel UInt8, neg UInt8, tgt_mis UInt8, raw UInt32)— names and types matchsqlcpu/schema.sqlliterally, as every other table in this section does.The seven flags after
sgare decode-time collapses. Each is a pure function of the instruction word and its address, so an execute expression may select a shared primitive with them instead of carrying one arm per opcode:m_sg1/m_sg2/m_hifor the four multiplies,d_sgfor the four divide and remainder forms,cmp_sel/negfor the six branches, andtgt_misfor the eager misaligned-target halt onjaland on a taken branch.tgt_misis meaningless forjalr, whose target is register-relative and whose alignment is only knowable at execution.idis the collapsed opcode space, including dedicated arms for the fatal-halt decode cases (§1):ecall,ebreak, CSR, and unimplemented/illegal each get their ownid, disjoint from the executable arms.immis already sign-extended;tgtholds only the absolute branch/jump target as a byte address — not a word index — (a word index discards bit 1, which is exactly the bit misaligned-target detection needs; an earlier draft was word-indexed and was reverted for that reason, see §1's eager alignment check) for branches andjal(notjalr, which is register-relative and computed live) — it is never the link value. The link valuejal/jalrwrite tord(pc + 4) is not stored indecoded; it is computed live from the executing pc, since it is simple pc-relative arithmetic with nothing to gain from precomputing. An earlier draft of this table usedtargetfor both the jump target and the link value on the same row — caught in review (issue discussion on PR #42/#48) before it shipped as a real bug: harmless in the Phase 0 benchmark this design is descended from, since that benchmark's decode data is synthetic and never executed (docs/experiments/arrayfold-baseline.md), but a genuine correctness bug in a table meant to produce correct results.rawcarries the original instruction word, needed only for theILLEGAL_INSNhalt record (§1) since every other column replaces the raw word rather than preserving it. Decoding must happen inside ClickHouse — decoding externally and inserting the result is a PURITY.md violation, not an optimization.
Materialize ram into the batch's constant array with FINAL, not
argMax(value, version) ... GROUP BY word_addr: measured 0.022–0.030 s against
0.245–0.256 s, and FINAL stayed flat with 1.2 M accumulated store deltas.
The driver invokes one batch = one INSERT ... SELECT executing up to K
instructions (K default 50,000; tunable).
The write-log high-water mark bounds K from above before cost does. A
boot-window batch stops on the mark after 60,006 retired instructions
whatever K says, and arrayFold runs every element of range(K) whether
or not it retires, so a larger K there buys iterations that retire nothing
(docs/experiments/batch-attribution.md). Over a fixed
120,000-instruction window, batch cost fits a fixed setup plus per-step work
plus superlinear write-log growth, which puts the optimum at K ≈ 47,900,
0.4% under the cost at K = 60,000 (same record). The default sits inside that
flat region. Per-batch fixed cost is 624 ms on ClickHouse 26.7.5.10,
4.8% of a 13,088 ms steady-state batch, and it is the analyzer walking the
generated SQL rather than the RAM capture (same record).
Measured real-ROM throughput is stated by the record docs/benchmarks.md
indexes as current for the pinned server, at K = 60,000 and HWM = 20,000,
chained batches past the compile threshold, one fresh container per arm, on a
quiet machine. A figure quoted here carries the window, K, the high-water
mark, the batch range and the server version, and moves when the pin moves.
A figure taken over the first three batches of a series reads 18.3%
lower, because those are both the uncompiled batches and, in boot, the
write-log-saturated ones
(docs/experiments/short-circuit-and-gameplay.md).
The bar to clear is ≥5,000 instructions/sec, stretch ≥11,000.
Against §1's measured -timedemo demo3 length those are a five-day run and a
two-and-a-half-day one.
A batch ends early on: halt,
FRAME_COMMIT write, or write-log high-water mark. Batch commit is atomic:
either all effects (ram deltas, cpu_state row, MMIO side effects) land or
none do. "Atomic" is a statement about externally observable state, not a
requirement for a single cross-table transaction — ClickHouse has none. A
design satisfies this contract if the batch's committed facts are captured
in one atomic single-row write (batch_commit, §5), and every other effect
converges to match that row through derivation that is safe to redo any
number of times, such that no observer — including a SPEC §7 differential
trace, before or after a crash and restart — can ever witness a
half-applied batch. The driver's only logic: loop batches, insert key
events, blit committed frames. See PURITY.md.
Both refemu and sqlcpu must emit identical checkpoints:
- Every
CHECKPOINT_INTERVAL(default 4,096) retired instructions:(icount, pc, xxh64(pc || regs[1..31] as LE bytes)). - Every
RAM_HASH_INTERVAL(default 1,048,576) instructions: additionallyxxh64over the full RAM region as LE words, and a second, independentxxh64over FRAMEBUFFER (64,000 B) concatenated with PALETTE (768 B), both in address-ascending order (64,768 bytes total) —fbhash. MMIO is excluded from both hashes: it is live device state, not a value two independently-running engines are expected to agree on bit-for-bit.fbhashexists because a store that lands the wrong value at the right framebuffer/palette address (or vice versa) touches neither a register nor RAM, soreghash/ramhashalone are blind to exactly the class of bug most likely to matter for DOOM specifically — a rendering bug that wouldn't surface until the final frame comparison, with no checkpoint in between narrowing down where it happened. A separate column (not folded intoramhash) trades a larger trace-format surface for telling a divergence hunt which region diverged without bisecting — real diagnostic value specifically in the Phase 3 desync hunt this format exists for. Agreed byrefemuandsqlcpuindependently (issue #55). - Trace file: one checkpoint per line, TSV:
icount<TAB>pc_hex<TAB>reghash_hex[<TAB>ramhash_hex<TAB>fbhash_hex]. - First divergence = first line that differs.
make diff N=<count>runs both engines N instructions and reports it; divergences are filed with thedivergence-reportissue form.
- No wall-clock, no randomness, no host environment reads on any computation path.
- Iteration order that affects results must be explicit (
ORDER BYeverywhere it matters; never rely on ClickHouse block order). - ClickHouse server version is pinned repo-wide (see
docker-compose.ymland workflow files — keep in sync). Bumps areci:PRs with full nightly deep-diff evidence. - The ROM build must be byte-reproducible (§4).
- Every query setting a computation depends on is pinned in that query's
own
SETTINGSclause, never inherited from a session, a user profile or a server default.short_circuit_function_evaluationis one such setting: whether an unselectediformultiIfarm evaluates decides whether a guarded division faults, so the fold names it rather than inheriting it. A setting that only affects speed does not belong here.
Resolved by the Phase 0 benchmark (evidence:
docs/experiments/arrayfold-baseline.md; decisions: ADR-0001, ADR-0002):
- arrayFold throughput.
arrayFoldcarries a CPU step, and pre-decoding is the lever (7.4× on the same fold, ADR-0002). ADR-0004 retired ADR-0001's ≥10,000 instructions/sec acceptance criterion as a merge gate for correctness work. Measured real-ROM throughput is whatdocs/benchmarks.mdindexes for the pinned server; §6 states the conditions. Whether the gameplay window still clears the ≥5,000 bar on 26.8.2.7 is open, and a quiet-machine run settles it. The architectural decision — arrayFold, write-log memory, K = 50,000 as the default — stands. - Accumulator copy with large captured constant arrays. Does not happen. Fold throughput is flat across a 6,144× range in captured-array size (4 KiB → 24 MiB: 113,895 vs 106,951 instructions/sec). Holding all 24 MiB of RAM as a query-level constant is sound.
- Fallback decision. Not needed — no fallback in ADR-0001's list was used. The decisive change was moving decode out of the fold lambda into a table (ADR-0002, 7.4× on the same fold), which the ADR did not anticipate because it assumed the cost model was about data movement rather than expression-node count.
-
Kdefault. 50,000, fixed in §6 above. - ClickHouse version pin. 26.8.2.7, pinned by image digest
(
sha256:fa394da8…, the multi-arch index) indocker-compose.yml, which every workflow reads throughmake up; bumped from 26.7.5.10 on #295. Emulation costs 0.88x on the gameplay window end to end and the resident simulation statement's analysis costs 3.07x less (docs/experiments/clickhouse-26-8.md). Note the image restricts thedefaultuser to container-local addresses; the pin shipsCLICKHOUSE_PASSWORDindocker-compose.ymland in every CI service container so host and runner connections work (see issue #3).
Deferred, with the milestone that closes them:
-
IPMSfinal value (§3.1). Deliberately not derived from measured throughput — elastic time exists precisely so that emulator speed cannot affect emulated behaviour. It is a game-speed parameter:IPMS= 10,000 means DOOM believes one millisecond passes per 10,000 retired instructions. Validated at ROM bring-up, whenrefemucan report how many instructions a real DOOM tic actually costs. Owner:refemu, Phase 1. - MMIO addresses and framebuffer format (§2, §3). Cannot be ratified
before the ROM boots; doomgeneric's platform layer may want a register
this table does not have. Owner:
rom, at the Phase 1 milestone.