Skip to content

Add multi-CPU-variant support (68010/CPU32/68020/68030/68040/68060) - #2

Merged
jenska merged 1 commit into
mainfrom
cpu-variant-support
Sep 12, 2026
Merged

jenska merged 1 commit into
mainfrom
cpu-variant-support

Conversation

@jenska

@jenska jenska commented Sep 12, 2026

Copy link
Copy Markdown
Owner

Summary

m68kdasm previously decoded plain 68000 opcodes only. This PR adds:

  • A CPU selection mechanism — DecodeOptions.CPU, defaulting to M68000 which preserves prior decode output byte-for-byte.
  • Per-opcode/addressing-mode gating by an explicit CPU bitset (not an ordinal comparison), since the 68k family isn't a strict newer-implies-older chain: CPU32 branches off 68010 with its own extensions, and 68040 drops CALLM/RTM that 68020/68030 have.
  • 68010/CPU32: MOVEC, MOVES, RTD, BGND.
  • 68020+: full extension-word addressing (memory indirect, scaled/suppressed index, 0/16/32-bit base and outer displacements), the BFxxx bitfield family, CAS, CHK2/CMP2, 32×32 MULU.L/MULS.L/DIVU.L/DIVS.L, PACK/UNPK, CALLM/RTM, TRAPcc, LINK.L, EXTB.L, CHK.L.
  • 68030/68040/68060 inherit the 68020 additions automatically via the CPU bitset tagging (68040/68060 correctly exclude CALLM/RTM).

Deliberately not implemented, documented in docs/design-cpu-variants.md: CAS2, CPU32's TBLS/TBLU family, 68040's MOVE16/CINV/CPUSH, and FPU/PMMU coprocessor instructions. Each carries enough encoding uncertainty or scope that a guessed implementation felt worse than a documented gap.

Bonus: plain-68000 baseline fixes

Auditing the opcode table for this work turned up several pre-existing gaps unrelated to CPU variants — some silent:

  • Missing entirely: LINK, UNLK, EXT, CHK, EXG, RESET, RTE, RTR, ILLEGAL, NBCD, MOVEP, and all of ADDQ/SUBQ/Scc/DBcc.
  • Silently mis-decoding as something else: TAS fell through to TST, ADDX/SUBX fell through to ADD/SUB with a garbled operand, EXG fell through to AND.

Every canonical 68000 mnemonic now has a decoder.

Reorganization

  • Split the overloaded types.go into types.go (data model), cpu.go (CPU/cpuSet), and opcodetable.go (the single opcode-dispatch table).
  • Folded four session-specific files back into the existing per-instruction-family files (move.go, arithmetic.go, compare.go, bcd.go, multiply_divide.go, single_op.go, special.go).
  • Deduplicated several near-identical decoder functions behind small shared factories (noOperand, singleRegisterOperand).

Notes for reviewers

  • go.sum was previously out of sync with go.mod (missing the v1.4.0 entry it required) — fixed via go mod tidy.
  • Found and reported a real regression in the m68kasm test dependency (PC-relative literal displacements stopped being adjusted for the extension-word base starting in v1.3.2, still present in v1.4.0) while chasing a flaky round-trip test; this repo's own decoder now consistently uses the raw encoded value for PC-relative display (matching the current m68kasm behavior), documented inline where relevant.
  • 65 tests across both packages, including a full-table regression test (TestFindDecoderMatchesOpcodeTable) that asserts no two opcode patterns in the dispatch table silently collide.

🤖 Generated with Claude Code

m68kdasm previously decoded plain 68000 opcodes only. This adds a CPU
selection mechanism (DecodeOptions.CPU, default M68000 preserves prior
behavior byte-for-byte) and gates every opcode/addressing-mode pattern
by an explicit per-CPU bitset, since the 68k family isn't a strict
newer-implies-older chain (CPU32 branches off 68010 with its own
extensions; 68040 drops CALLM/RTM that 68020/68030 have).

New capability by CPU:
- 68010/CPU32: MOVEC, MOVES, RTD, BGND
- 68020+: full extension-word addressing (memory indirect, scaled/
  suppressed index, 0/16/32-bit base and outer displacements), BFxxx
  bitfield family, CAS, CHK2/CMP2, 32x32 MULU.L/MULS.L/DIVU.L/DIVS.L,
  PACK/UNPK, CALLM/RTM, TRAPcc, LINK.L, EXTB.L, CHK.L
- 68030/68040/68060 inherit the 68020 additions automatically via the
  CPU bitset tagging (68040/68060 correctly exclude CALLM/RTM)

Deliberately not implemented (documented in docs/design-cpu-variants.md):
CAS2, CPU32's TBLS/TBLU family, 68040's MOVE16/CINV/CPUSH, and FPU/PMMU
coprocessor instructions — each carries enough encoding uncertainty or
scope to warrant its own follow-up rather than a guessed implementation.

While auditing the opcode table for this work, also found and fixed
several pre-existing plain-68000 gaps unrelated to CPU variants: LINK,
UNLK, EXT, CHK, EXG, RESET, RTE, RTR, ILLEGAL, NBCD, MOVEP, and all of
ADDQ/SUBQ/Scc/DBcc were missing entirely, and TAS/ADDX/SUBX/EXG were
silently mis-decoding as other instructions (e.g. EXG as AND, TAS as
TST) due to opcode-space collisions the original table didn't account
for. Every canonical 68000 mnemonic now has a decoder.

Also reorganized internal/decoders/*.go: split the overloaded types.go
into types.go (data model), cpu.go (CPU/cpuSet), and opcodetable.go
(the single opcode-dispatch table), folded four session-specific files
back into the existing per-instruction-family files, and deduplicated
several near-identical decoder functions behind small shared factories.

Covered by 65 new/updated tests across the decoder package and the
top-level package, including a full-table regression test asserting
no two opcode patterns silently collide.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@jenska
jenska merged commit 5c40fc0 into main Sep 12, 2026
1 check passed
@jenska
jenska deleted the cpu-variant-support branch September 12, 2026 20:06

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 8e89194c30

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

valSUBXMemW = 0x9148
valSUBXMemL = 0x9188

valBFTST = 0xE0C0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Register bitfield opcodes at E8C0–EFC0

The bitfield opcode family is offset by 0x0800: the design document identifies it as 0xE8C0–0xEFC0, while these constants register 0xE0C0–0xE7C0. Consequently, real 68020+ bitfield instructions fall through to the shift/rotate decoder, while matching memory shift encodings are claimed as BFxxx instructions and incorrectly consume a bitfield extension word, desynchronizing range disassembly.

Useful? React with 👍 / 👎.

Comment on lines +57 to +60
if ext&0x8000 != 0 {
mnemonic = mnemonicSigned
}
dlMeta := registerOperand(RegisterKindData, uint8(ext&0x7))

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Decode long MUL/DIV extension fields correctly

The long multiply/divide extension word does not use bit 15 for signedness or its low three bits as the sole 32-bit destination register. For example, the valid MULS.L D0,D2 encoding 4C 00 28 00 is decoded here as MULU.L D0,D0; unsigned operations targeting any register other than D0 are likewise rendered with D0. This shared parser also misdecodes the corresponding DIV.L forms.

Useful? React with 👍 / 👎.

Comment on lines +117 to +118
dcMeta := registerOperand(RegisterKindData, uint8((ext>>3)&0x7))
duMeta := registerOperand(RegisterKindData, uint8(ext&0x7))

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Extract CAS register operands from their encoded fields

CAS stores its compare and update data-register numbers in different extension-word fields than these shifts use. For the valid CAS.W D1,D2,(A3) bytes 0C D3 00 81, these lines produce CAS.W D0,D1,(A3) instead of preserving D1 and D2, so nearly every nonzero CAS register combination is reported incorrectly.

Useful? React with 👍 / 👎.

Comment on lines +90 to +93
disp := int32(int16(binary.BigEndian.Uint16(data[2:4])))
// Matches decodeBxx's existing displacement-base convention for the
// 16-bit form (offset counted after consuming the displacement word).
target := uint32(int32(inst.Address) + 4 + disp)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Base DBcc targets on the extension-word PC

DBcc's signed displacement is relative to the PC immediately after the opcode word, not after its displacement extension. Thus DBF D0,* (51 C8 FF FE) at $1000 should target $1000, but this calculation returns $1002; the incorrect target also propagates into branch metadata and symbolization for every DBcc instruction.

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant