Boot progression (4.3M instructions total):
┌──────────┬──────────┬──────────────────────────────────────────────────────────────────────────────┐ │ Line │ PC Range │ Description │ ├──────────┼──────────┼──────────────────────────────────────────────────────────────────────────────┤ │ 1 │ $000000 │ Reset vector │ ├──────────┼──────────┼──────────────────────────────────────────────────────────────────────────────┤ │ 568 │ $A02Exx │ Early ROM init │ ├──────────┼──────────┼──────────────────────────────────────────────────────────────────────────────┤ │ 2190 │ $A463xx │ Main ROM startup │ ├──────────┼──────────┼──────────────────────────────────────────────────────────────────────────────┤ │ 4195 │ $A14Cxx │ Hardware init routines │ ├──────────┼──────────┼──────────────────────────────────────────────────────────────────────────────┤ │ 31909 │ $A464DE │ RAM test/sizing (the $A46AF0 checksum loop runs from here until line ~1.34M) │ ├──────────┼──────────┼──────────────────────────────────────────────────────────────────────────────┤ │ 1342662 │ $A02Fxx │ Post-RAM-test │ ├──────────┼──────────┼──────────────────────────────────────────────────────────────────────────────┤ │ 1350043 │ Various │ ROM initialization continues ($A465xx, $A4A5xx) │ ├──────────┼──────────┼──────────────────────────────────────────────────────────────────────────────┤ │ 1352294 │ $A466DA │ Enters memory clear/init region │ ├──────────┼──────────┼──────────────────────────────────────────────────────────────────────────────┤ │ 1352353+ │ $A4685E │ Memory init loop with MOVEM/EOR pattern (part of the RAM march) │ └──────────┴──────────┴──────────────────────────────────────────────────────────────────────────────┘
The CPU advances through reset, early init, RAM testing, and into memory initialization. The $A4685E-$A46880 MOVEM loop above is part of the legitimate RAM march, not the real blocker. See the investigation log below for where the boot actually stalls and why video never initialises.
Goal: boot now runs but the screen stays black. Used MAME (maclc, same ROM) as ground
truth alongside the Verilator sim.
- FIXED & committed (
2206dfb): address-decoder motherboard-RAM mapping. Made the RAM march terminate and advanced boot to the bank-scan region (the same region MAME reaches). - VRAM→DTACK done (uncommitted, tree change): VRAM (
$F40000-$FBFFFF) was on the 6800 VPA path; now routed to async DTACK like RAM, in BOTHverilator/sim.vANDMacLC.sv. This is architecturally correct (matches MAME and the lbmactwo NuBus-on-DTACK core) and makes the VRAM data path bit-identical to RAM — but it is NOT the lever: screen still black, probe still fails at the same point. Keep it; it's a prerequisite, not the fix. - ACTUAL ROOT CAUSE (2026-06-02 — CONFIRMED via execOPC/Flags probe; corrects the earlier
"stale operand" wording in section 6): a TG68K flag-commit-vs-branch pipeline race. The
bank probe does two consecutive
cmp.llongword reads — aliascmp.l (A0,D2)@$A467ECthen realcmp.l (A0)@$A467F2. At the real cmp'sexecOPCcommit BOTH ALU operands are correct (OP1=D0=5368656C,OP2=mem=5368656C) so it computes EQUAL and latchesZ=1— but ONE CYCLE LATE. The followingbeq $A467F4already evaluated its condition using the staleZ=0from the alias cmp → falls through to$A467F6(bank absent) → probe fails.cmp.l (An)uses the tightget_ea_nowimmediate-read path (2-cycle longword read delays the flag commit); the workingcmp.l (d16,An)defers viald_dAn1, whose extra cycle aligns the commit ahead of the branch. NOT VRAM-specific (reproduces on VPA and DTACK); the data path is 100% clean. Fix = give the(An)longword op the same one-cycle alignment (mirrorld_dAn1), or stall the Bcc on a pending flag commit. (VHDL is now the source of truth — see the consolidation commit.) - FIX IMPLEMENTED & VERIFIED 2026-06-02 (uncommitted): added micro_state
ld_An1(TG68K_Pack.vhd/.sv) and split the(An)source-decode inTG68KdotC_Kernel.vhd— for longword plain(An)(opcode(5:3)="010" AND datatype="10") it now doessetstate<="01"; next_micro_state<=ld_An1(one INTERNAL cycle, no code fetch — plain (An) has no extension word), andld_An1does the usualget_ea_now+setnextpass.(An)+/-(An)and byte/word(An)are untouched. Regen viaconvert_to_verilog.sh. Verified with temp kernel debug ports (execOPC/Z/exe_condition + bra1/ld_An1 taps, since removed): thebeq $A467F4now samplesZ=1/cond=1and branches to$A467FE(bank present); the clean build's CPU trace then advances throughjmp (A6)to the bank-scan driver$A4A590($A4A6xx) — the same post-probe region MAME reaches. The frame-350 screenshot is STILL black, but ONLY because the correct RAM-test march ($A468xx, now testing real RAM at$000000per2206dfb) is long and not finished by frame 350 in sim — NOT a regression (see section 7). NOTE: the "$A467F6count" grep is a BOGUS oracle (it's the prefetch word after the beq, present on pass AND fail paths); use the beq's branch target / bank-scan reach instead.
This sim has clean-vs-incremental build nondeterminism. Several "it works!" results were
incremental builds with stale objects = FALSE POSITIVES. ALWAYS make clean && make
before trusting any result. (Also bit us at session start: an incremental build hung at
$A14E60; the clean build ran fine.)
- 68k ROM
releases/boot0.romis byte-identical to MAME350eacf0.rom(SHA16bef5853...). - The Egret HC05 firmware the sim actually loads (
rtl/egret/egret_rom.hex, SHA18b0dae3...) == MAME's default BIOS341s0851. - Cleanup nit: the unused
rtl/egret_rom.bin/rtl/egret_rom.hexin the REPO ROOT are the older341s0850— harmless but worth deleting to avoid confusion.
Identical ROM + 2 MB boots to video (grey desktop, mouse, blinking "?" disk). Its RAM march
runs ~2 passes then exits and never reads the $200000-$7FFFFF gap (debugger
watchpoints: 0 hits). Source: ~/repos/mame/src/mame/apple/{maclc,v8}.cpp. Romset:
/private/tmp/goodroms/maclc. Debugger gotcha: breakpoints/trace default to the Egret
HC05 — target maincpu explicitly (trace file,maincpu). 68k ROM runs at $00A4xxxx;
address mask is global_mask(0x80ffffff).
rtl/addrDecoder.v only asserted selectRAM for the motherboard mirror at $800000-$9FFFFF.
With no SIMM (ram_config=0x24), $000000-$1FFFFF fell through to selectUnmapped (returns
$FFFF). MAME's v8.cpp ram_size() ALWAYS installs the motherboard RAM at mb_location
(= SIMM size = 0 when no SIMM) plus the $800000 mirror. Added in_motherboard_low. Result:
the RAM march terminates instead of grinding forever, and boot reaches the post-march
bank-scan ($A4A6xx). (The unmapped-read VALUE, $FFFF vs MAME's $0000, was tested and is
NOT the lever — leave it $FFFF.)
Found by PC-stream diffing our trace vs a MAME maincpu trace, anchored at RAM-sizing entry
$A467CC. First divergence is the bank-presence probe:
A467E6 move.l D0,(A0) ; write pattern $5368656C to bank base A0
A467EC cmp.l (A0,D2.l),D0 ; D2=$40000 -> read $F80000 (alias check)
A467F0 beq $a46804 ; if aliased
A467F2 cmp.l (A0),D0 ; read back $F40000
A467F4 beq $a467fe ; MAME: TAKEN (readback OK) OURS: NOT taken (readback wrong)
A467F6 bclr #0,(A2) ; ours-only: bank marked "not present"
A0 = $50F40000 = the VRAM aperture. The ROM probe table at ROM offset $3B00 holds the
I/O/video bank bases ($50F40000=VRAM, $50F26000=PseudoVIA, $50F0xxxx=VIA...). MAME reads
the pattern back; we do not -> the video bank is mis-marked -> that corrupts the RAM-config
descriptor table -> the later march gets a garbage base ($4E754F00, in the unmapped gap) and
wanders. This is the black-screen root cause and it is on the video memory path.
Why VPA: $F40000 has cpuAddr[23:21]==3'b111, which the CPU bus logic treats as a 6800
E-clock VPA peripheral:
sim.v:262-263 / MacLC.sv:845-846
_cpuVPA = (cpuFC==3'b111) ? 0 : ~(!_cpuAS && cpuAddr[23:21]==3'b111);
_cpuDTACK = ~(!_cpuAS && cpuAddr[23:21]!=3'b111) | !dtack_en;
VRAM is SDRAM-backed and must use async DTACK like RAM. On the VPA path the CPU samples the read on a fixed E-clock phase that catches the SDRAM read BEFORE it settles -> stale data.
Instrumented sim_ram.v + sim.v with an armed cycle-by-cycle $display of
busCycle/videoBusControl/memoryLatch/selectVRAM/memoryAddr/ram_do/dataControllerDataOut/
AS/DTACK/cpuAddr/busstate, plus clk16-enable sampling and the TG68 wrapper's internal
tg68k.tg68_din_r / s_state. Findings (all CLEAN-build):
- The data path is fine:
dataControllerDataOutsettles to the correct word at the first cpu-slotmemoryLatchand is rock-stable for 40+ clocks (cpu_data only latches in cpu slots, so interleaved video fetches don't corrupt it). sim_ramhas a 1-clock REGISTERED-READ lag (dout<=mem[addr]): whenmemoryLatchsamples the cycle right aftermemoryAddrswitches,ram_dois still the prior address.- The longword WRITE was being TRUNCATED: with VRAM on DTACK,
mem[$580001]was written 0 times — DTACK acked before the 2nd word's cpu write slot, so$656Cwas LOST and the readback legitimately returned 0. - Sub-fixes that DID work in a clean build: (a) combinational
sim_ramread removes the lag; (b) avram_slotcounter giving each VRAM access TWO cpu slots (1st settles read / commits write, 2nd asserts DTACK). After these:mem[$580001]IS written, and the dump shows the TG68's OWN input registertg68_din_rlatch$5368(word1) then$656C(word2) at state 6/7 — i.e. the CPU input register holds exactly$5368656C=D0. - The old "paradox" — NOW RESOLVED (2026-06-02), see section 6.
tg68_din_ris correct for both words yetcmp.lreports unequal. The discrepancy is inside the TG68 kernel, exactly as suspected — a back-to-back-longword-read operand bug, NOT a VRAM/bus issue.
Method: re-ran (VRAM on DTACK) with three ifdef SIMULATION probes — sim_ram VRAM-range
read/write log, the wrapper's tg68_din_r latch + per-clkena (s_state/busstate/din_r)
log, and an in-kernel dump of tg68_pc/data_in/data_read/last_data_in/last_data_read/
memmask/state (the readable net aliases exist in the generated .v). All probes removed
afterward; tree holds only the DTACK change.
The probe code at $A467E6 is a write-then-read-back bank presence test:
A467E6 move.l D0,(A0) ; D0=$5368656C -> write to $50F40000 (VRAM base)
A467EC cmp.l (A0,D2.l),D0 ; D2=$40000 -> read $50F80000 (alias) = $00000000; not-equal (correct)
A467F0 beq $a46804 ; not taken (alias != D0, good)
A467F2 cmp.l (A0),D0 ; read $50F40000 = $5368656C; SHOULD be equal to D0=$5368656C
A467F4 beq $a467fe ; SHOULD take, but does NOT -> bank wrongly marked absent
What the kernel dump proves — operand assembly is CORRECT for BOTH the failing and a known-good
cmp.l, so the assembly is not the bug:
FAILING cmp.l (A0) @A467F2: reads 5368 then 656c -> data_read=5368656c, last_data_in=5368656c OK
WORKING cmp.l ($24,A1) @A02F90: reads 5600 then 0000 -> data_read=56000000, last_data_in=56000000 OK (bne falls through = equal)
Both deliver the right 32-bit operand into last_data_in. The distinguishing factor: the FAILING
cmp.l (A0) is the second of two back-to-back longword reads — preceded immediately by the
alias cmp.l (A0,D2) @ $A467EC whose operand is $00000000. The working one is a lone read.
D0 is provably correct ($5368656C): the alias cmp @ $A467EC returned not-equal, which it
could only do if D0 != $00000000. So with operand=$5368656C AND D0=$5368656C both correct,
the second cmp still reporting not-equal means the ALU received a STALE operand — the alias's
$00000000 carried over from $A467EC instead of the freshly-read $5368656C. The kernel does
not flush/advance the compare operand between two consecutive longword reads to the same base reg.
- This is addressing-mode / instruction-sequence general, NOT VRAM-specific. It reproduces on
VPA and on DTACK; memory/sim_ram/wrapper/kernel
data_readare all correct. The DTACK change is correct but irrelevant to this bug. last_data_readis a red herring: it tracks opcode fetches (b090/6708/08aa...), not the data operand, in BOTH the working and failing cases — so it is NOT the ALU operand source.
(superseded scratch note below; kept for the cycle timestamps)
The kernel latches last_data_read only when state="00" OR exec(update_ld)
(TG68KdotC_Kernel.vhd:495). For this read the assembled longword is valid at state="10"
(@760); by the time state="00" comes (@768) an interleaved opcode prefetch has already
overwritten data_read. So cmp.l compares D0 against $00006708 (stale) → not equal →
$A467F4 beq not taken → $A467F6 marks the bank "not present" → descriptor table corrupts →
$4E754F00/$4F.. gap-march → no video.
- The alias probe
cmp.l (A0,D2),D0@$A467EChits the SAME mis-capture, but its operand is$00000000, so "not equal" is coincidentally correct → the bug is masked there. - Not a kernel-version issue: lbmactwo's NEWER kernel has the IDENTICAL capture condition
(
TG68KdotC_Kernel.vhd:537) and the IDENTICAL longword-assembly process, yet boots to video. So the difference is in bus/wrapper TIMING (when the prefetch interleaves vs the operand read), not the kernel's capture logic. The MacLCtg68k.vwrapper differs from lbmactwo's. - Not VPA/DTACK and not the data path: memory, sim_ram,
cpu_data,tg68_din_r, and the kernel's owndata_readare all correct. The previous "VPA stale data" theory is wrong; the DTACK change is correct but insufficient.
- Compare the failing operand read against a WORKING longword cmp/read on RAM, capturing
the kernel
state/data_read/last_data_read/exec(update_ld)sequence for both. The question to answer: in the working case doesexec(update_ld)fire atstate="10"(capturing the operand before the prefetch), or does the prefetch simply not interleave there? That tells you whether the fix belongs in (a) decode (update_ldforcmp.l (An)), or (b) the wrapper's fetch/read sequencing (skipFetch/busstate handling). - Most promising fix: port the lbmactwo
tg68k.vwrapper (and its matching kernel). lbmactwo is a working core that does longword reads from DTACK video memory with the same kernel logic; its wrapper sequences fetch-vs-read differently. This needs the newer kernel's extra ports (longword,VBR_out,cpu_halted,berr_inhibit,berr_data,IPL_autovector=1). Big, risky change (CLAUDE.md warns on CPU edits) — do it in a worktree and gate on the frame-350 screenshot + the verification below. Bonus: would also bring real BERR-on-unmapped support. - Verify any fix, in a CLEAN build:
$A467F6count -> 0 (probe passes),$4E754F00/$4F..march gone, framebuffer writes to$F40000appear, screenshot goes black -> grey/Happy-Mac. Cross-check V8 video registers vsmonitor_id/video_modedefaults (monitor_id=2 / 512x384, video_mode=2 / 4bpp). - Separately confirm (or rule out) the post-fix
$4F62xxxxgap-march as legit HMMU-truncation vs another decode divergence — MAME never reads that gap.
master (= merge-base bc7019c) shows the grey memory-test PATTERN at frame 350 (PNG ~16.6KB);
this branch (HEAD) shows BLACK (~9KB). This is the EXPECTED consequence of the 2206dfb
RAM-mapping correctness fix, not a video-path regression:
- master does NOT contain
2206dfb, so its$000000-$1FFFFFreturns$FFFFopen-bus garbage. The RAM-test march ($A468xx:eor.l D2,(A2)/move.l (A2)+,D2/dbf %d3@$A46904) reads that garbage, its checksum/compare loop exits EARLY, and it lands in$A469xxdoing garbage ops that happen to touch$50F4xxxx(VRAM) -> the visible "pattern." - HEAD has
2206dfb-> real RAM at$000000-> the march runs its FULL, CORRECT, multi-pass length (more RAM to test) -> still mid-march at frame 350 -> black. Confirmed HEAD PROGRESSES (a frame-700 run reached the post-march caller$A46AF8then looped back for another pass — not hung, just slow in sim; frame 700 ~= cycle 560M). So master's pattern was an ARTIFACT of the pre-fix RAM bug; HEAD's black is the correct long-march-still-running state. Scanout works (master proves the path). Seeing real video in SIM needs a much longer headless run (or just verify on FPGA, which runs the march in real time). Do NOT revert2206dfbor5c3c94f. The Row-M orange overlay difference is a separate red herring (a HEAD-only debug row).
The committed cmp.l fix (42ae7a6) makes the bank probe pass; boot now reaches the real
RAM-test march. The sim screen shows a whole lot of orange (orange/grey vertical stripes,
not black) — this is uninitialized VRAM being scanned out through the 4bpp palette. It PROVES
the V8 scanout path works end-to-end; the ROM just hasn't written a real framebuffer yet because
it is still in the march. (master's "grey pattern" is the same phenomenon, different palette.)
The march is a ~21-pass moving-inversion test, not 2 passes. $A46910 cmpi.w #21,%d7; beq $A4694C gates up to 21 iterations; $A46918 cmp.l d0,d1 evolves the pattern and loops back to
the inner march at $A468a8. MAME runs all ~21 passes but its 68020 instruction cache holds
the entire unrolled move.l (A2)+,D2 / eor.l D2,(A2) loop, so it rips through them in ~127 MAME
frames (per-frame PC sampler: march at MAME F97-164 + F195-222, bank-scan $A4A4xx at F188, late
init by F231). Our cacheless TG68 re-fetches every opcode AND is bus-starved by video DMA, so
each pass costs ~150-280 sim frames.
Measured sim cadence (clean build, monitor_id=2/512x384, --no-cpu-trace): pass-boundary
$A46910 hits at F110, F267, F548, F712, F876 (Δ ≈ 157, 281, 164, 164 → ~191 frames/pass).
Extrapolated march completion ($A4694C) ≈ F3900-4000, after which the ROM video-init writes
the real framebuffer at $F40000→SDRAM $580000. So seeing real video in SIM needs a run to
~F4000+; the FPGA runs the same march in real time (a multi-second RAM-test pause). This is NOT a
bug and NOT a regression — it is the faithful, long, correct march. Do NOT shrink RAM or revert
RAM/VRAM commits to "fix" it (RAM-shrink risks diverging from MAME's detected size/descriptors).
Instrumentation added this session (verilator/sim_main.cpp): --no-cpu-trace flag (skips the
per-instruction Musashi disasm/log that ballooned to 300MB and dominated runtime — required for
long runs); 5M-cycle periodic PC+frame sampler; and march-progress counters that print on
$A46910/$A4694C/$A4A590. Sim monitor_id changed 6→2 (sim.v:480) to match the FPGA
default (status[11:10]=0 → 512x384) AND cut video-DMA contention (~1.5x faster march). Note
$A4A590 and $A46AFx are shared routines hit in EARLY boot too — only $A4694C is a reliable
march-done marker.
Section 8's "faithful, long, correct march" conclusion is WRONG — corrected here.
Measured MAME ground truth (Lua read-tap on maincpu program space, no debugger needed —
/tmp/tap910.lua, /tmp/tapregs.lua): MAME hits the shared region-test check $A46910
exactly 4 times for the whole boot (I/O, I/O, the 2MB RAM region, VRAM). Our core hits it
26+ and was still going — ~6x+ the memory-test work. THAT extra work, not cache or bus, is the
bulk of the ~60x slowdown (sim ~120s emulated vs MAME ~2s; real HW ≤30s for 10MB → ~6s for 2MB).
Root cause found by dumping the TG68 register file at each $A46910
(VERTOPINTERN->emu__DOT__tg68k__DOT__tg68k__DOT__regfile[0..15], D0-D7=0..7, A0-A7=8..15):
- MAME PASS#3 (the real 2MB RAM test):
A0=00800000 A1=009FFFECD7=3 D0=6DB6DB6D D1=B6DB6DB6. - OUR PASS#1: same D0/D1/D7 (
6DB6DB6D/B6DB6DB6/3) butA0=4E754EFA A1=350EACF0— ROM bytes ($350EACF0= ROM checksum/ID,$4E75= rts), and A0 > A1 (inverted). So our core reaches the right logical step with garbage region bounds → marches a corrupt, inverted range → loops nonsensically → slow AND no real RAM tested → no video. This IS the "$4E754F00gap-march" from sections 5-6.
IMPORTANT distinction: the cmp.l video-bank probe fix (42ae7a6) DOES work — instrumented
$A467FE (present) and $A467F6 (absent) both fire correctly at F45 (present for VRAM, absent for
a genuinely-absent bank). The garbage A0/A1 is a SEPARATE, still-present bug in the bank-sizing
/ region-descriptor build — NOT the video probe.
Bus-arbitration measurement (instrumented _cpuAS/_cpuDTACK taps in sim.v → sim_main per-clk
classify): inner march = ~9.6 clk_sys/access, stall ~30% (waiting for bus slot, the
arbitration-fixable part), xfer ~33%, idle ~36% (TG68 internal). So bus arbitration is worth only
~1.4x — MINOR vs the 6x descriptor bug. Deprioritized.
NEXT (resume here): trace where A0/A1 get loaded for the RAM region test — the code that should
produce $800000/$9FFFEC (MAME) — and find what corrupts it to $4E754EFA/$350EACF0. That is
the real fix for BOTH boot speed and video. Likely a corrupt region/descriptor table in low RAM or
a wrong pointer; cross-check against MAME's path between its PASS#2 ($A46910 D7=1) and PASS#3 (D7=3).
Sim instrumentation to reuse: [MARCH] PASS# regdump, [PROBE] present/absent, [BUS] stall
profile, --no-cpu-trace, 5M-cycle PC sampler (all in verilator/sim_main.cpp); MAME taps in
/tmp/tapregs.lua. Clocks are CORRECT (clk_sys 32.5MHz, CPU 16.25MHz — pll.v ×13/÷20); not a
clock or "running too fast" issue.
- Disasm:
docs/MacLC_ROM_disasm.txt(VMA 0x40800000; runtime$A0xxxx<-> disasm$4080xxxx, low 20 bits match). Probe$A467E6-$A467F6; RAM-sizing$A467CC; march$A468C4-$A46908; bank-scan driver$A4A600-$A4A6xx. - VRAM map: CPU
$F40000-$FBFFFF-> SDRAM word$580000(addrController_top.v:170-176). - CPU bus/VPA/DTACK:
sim.v:249-263,MacLC.sv:832-846; TG68 din sampletg68k.v:104-118. - MAME ground-truth how-to + fuller notes: auto-memory
video-march-investigation-2026-06-01andmame-ground-truth-maclc.