Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
23 commits
Select commit Hold shift + click to select a range
4f3d2fb
fix(inspect, od): render a node that has no byte position
hed0rah Sep 28, 2026
1ca72af
fix(au, aiff): two walkers found wrong by McGill's public test files
hed0rah Sep 26, 2026
9621395
fix(pdx): accept a slot table shorter than one bank
hed0rah Sep 26, 2026
006da09
fix(mdx): a cut module is still an MDX
hed0rah Sep 26, 2026
85afe65
fix(aiff, wav): a 12-bit sample is two bytes
hed0rah Sep 26, 2026
ab9cebc
fix(brstm): a cut stream says it is cut
hed0rah Sep 27, 2026
b197155
fix(au, aiff): the backported overrun warnings are plain strings
hed0rah Sep 28, 2026
1124520
fix(sigmf): the meta half places only its own bytes
hed0rah Sep 27, 2026
c285a63
fix(mp4): repair declines a multi-track file without raising
hed0rah Sep 27, 2026
faccd06
fix(mp4): render box timestamps without utcfromtimestamp
hed0rah Sep 27, 2026
87557a6
fix(audit): name an extract-only format and point at extract
hed0rah Sep 28, 2026
f74db47
fix(anomalies): a file that is a zip is not an appended zip
hed0rah Sep 28, 2026
ecbb7cc
fix(amxd): the marker is the device type
hed0rah Sep 28, 2026
a99276a
fix(vital): a skin is not a preset
hed0rah Sep 28, 2026
040e083
fix(wav): read the RF64 reservation instead of calling it damage
hed0rah Sep 28, 2026
80757fa
fix(wav): appended bytes past the RIFF end are not chunks
hed0rah Sep 28, 2026
c656f44
fix(sigmf, vital): the backported fixes use 1.8's paths
hed0rah Sep 28, 2026
3dec2d9
Release 1.8.7
hed0rah Sep 29, 2026
2645671
test: a repair-guard refusal is an answer, not a crash
hed0rah Sep 26, 2026
8cce5be
test(geometry): a format walk_file refuses is not a crash
hed0rah Sep 27, 2026
aff3c6d
fix(check): a valid RF64 is not a broken RIFF; a lost audio chunk is …
hed0rah Sep 29, 2026
69d9834
docs(changelog): 1.8.7 gains the RF64 and lost-audio check fixes
hed0rah Sep 29, 2026
5391e9b
fix(check): an overrunning audio chunk keeps its own explanation
hed0rah Sep 29, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ see exactly what a file is, flag anomalies, and edit or repair its structure.
Closer to readelf / 010 Editor / radare2's format layer than to exiftool, with
some optional audio analysis (BPM/key via librosa).

v1.8.6 · ~56k source LOC · ~52k test LOC · one hard dependency (`mutagen`);
v1.8.7 · ~56k source LOC · ~52k test LOC · one hard dependency (`mutagen`);
everything heavier is an optional, lazily imported extra, so `import acidcat`
pulls only the stdlib core.

Expand Down
51 changes: 51 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,57 @@ All notable changes to acidcat. The format follows
[Keep a Changelog](https://keepachangelog.com/en/1.1.0/); the project will
adopt [Semantic Versioning](https://semver.org/spec/v2.0.0.html) at 1.0.

## [1.8.7] - 2026-09-28

Bug fixes only, found by running acidcat over a curated library of real files
from every format it reads. No new features and no changes in behaviour
beyond the fixes.

### Fixed

- **`acidcat inspect` crashed on every SoundFont**, and so did `inspect
--full`, `--json` and `od`: the preset and instrument tree has no byte
position, and the renderers formatted its offset as a number. A test now
runs every render mode over every seed format.
- **Every valid RF64 failed `validate`.** RF64 writes 0xFFFFFFFF in its
32-bit size fields and keeps the real sizes in ds64; the checker read the
placeholders literally, reported the data chunk as overrunning the file,
and proposed an 88-byte RIFF size (repair refused to write it). RF64 is now
declined with a note, as `write` already declines it.
- **`validate` advertised a fix `repair` would refuse** on a file whose audio
chunk the walk cannot reach (a wrong size earlier in the file). It now
explains the misread and marks nothing repairable.
- **A WAV's RF64 reservation was reported as damage.** The 28-byte JUNK a
writer puts first is filled on purpose: with the ds64 sizes (sometimes
stale, the file having grown), with a quote Ableton Live writes there, or
with the RIFF size over such a quote. It is now read as the reservation,
and the quote names Ableton Live as the writer in `audit`.
- **Bytes appended past a WAV's RIFF end were read as chunks.** 2 MB of
appended zeros became 262,144 empty chunks and opening the file took
seconds. The walk now stops at the first id that is not four ASCII
characters; the bytes are still reported as trailing data.
- **Zip-based formats were reported as polyglots** (.xpn, .labx,
.multisample): the archive's own end record read as an appended zip, its
members as smuggled files. An SF3's Ogg samples were flagged the same way.
- **A Max for Live instrument was a "magic mismatch".** The four bytes at 8
are the device type: `aaaa` for an audio effect, `iiii` for an instrument.
- **A `.vitalskin` was read as a Vital preset**, every theme key flagged as
an unknown one.
- **A cut BRSTM was reported as whole.** The header's file size at 0x08 is
now read, and a file shorter than it declares says the stream is cut.
- **A SigMF recording opened by its `.sigmf-meta` was reported as broken.**
- **Every 12-bit AIFF and WAV was reported as damaged**: the size checks
rounded 12 bits down to one byte per sample.
- **Two walkers found wrong by McGill's public test files**: an AU whose
data runs past the end of the file, and an AIFF APPL chunk without a
fitting name.
- **76 X68000 PDX banks with a short slot table, and 16 cut MDX modules,
were not recognised.**
- **`constraints.repair` raised on a multi-track MP4** instead of declining
it, and MP4 timestamps used a call Python 3.12 deprecates.
- **`audit` called a console ROM `[unknown]`**; it now names the ROM and
suggests `acidcat extract`.

## [1.8.6] - 2026-09-25

### Fixed
Expand Down
6 changes: 3 additions & 3 deletions docs/formats/ableton-anatomy.html
Original file line number Diff line number Diff line change
Expand Up @@ -338,12 +338,12 @@
<!-- ===================== MAX FOR LIVE ===================== -->
<div class="panel" data-panel="m4l">
<div class="pfacts"><span>magic<b>"ampf"</b></span><span>header<b>12 bytes</b></span><span>chain<b>id + u32 length</b></span><span>payload<b>Max patcher JSON</b></span></div>
<p class="plede">A Max for Live device is an <b>Ableton Max Patch Format</b> container: the ASCII magic <code>ampf</code>, a version word, and then -- before the chunk chain begins -- a <b>four-byte marker</b>, <code>aaaa</code>. Only at offset 12 does the chain start: four-byte id, little-endian length, payload. The <code>ptch</code> chunk carries the Max patcher, a short binary preamble followed by JSON.</p>
<p class="plede">A Max for Live device is an <b>Ableton Max Patch Format</b> container: the ASCII magic <code>ampf</code>, a version word, and then -- before the chunk chain begins -- a <b>four-byte device type</b>: <code>aaaa</code> for an audio effect, <code>iiii</code> for an instrument. Only at offset 12 does the chain start: four-byte id, little-endian length, payload. The <code>ptch</code> chunk carries the Max patcher, a short binary preamble followed by JSON.</p>
<div class="sec">layout</div>
<pre class="code">0x00 "ampf" u32:version "aaaa" &lt;- 12-byte header
<pre class="code">0x00 "ampf" u32:version "aaaa" &lt;- device type; 12-byte header
0x0C "meta" u32:len ...
0x18 "ptch" u32:len mx@c ... { patcher JSON }</pre>
<div class="callout"><b>The marker is not a chunk, and mistaking it for one is loud.</b> Reading <code>aaaa</code> as an id makes the following four bytes -- the ASCII <code>meta</code> -- decode as a length of 1,635,018,093 bytes, and the whole chain collapses to one impossible chunk. The chain starts at <b>12</b>, not 8.</div>
<div class="callout"><b>The device type is not a chunk, and mistaking it for one is loud.</b> Reading <code>aaaa</code> as an id makes the following four bytes -- the ASCII <code>meta</code> -- decode as a length of 1,635,018,093 bytes, and the whole chain collapses to one impossible chunk. The chain starts at <b>12</b>, not 8.</div>
<div class="callout"><b>The length field cannot be trusted.</b> A chunk claiming more bytes than remain must stop the walk rather than read past the end. The <code>ptch</code> chunk ends exactly on the last byte of the file, so a chain that stops short means something is missing.</div>
</div>

Expand Down
2 changes: 1 addition & 1 deletion pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"

[project]
name = "acidcat"
version = "1.8.6"
version = "1.8.7"
description = "Binary analysis and reverse-engineering for audio file formats and music-production hardware"
readme = {file = "README.md", content-type = "text/markdown"}
license = {text = "MIT"}
Expand Down
2 changes: 1 addition & 1 deletion src/acidcat/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -46,7 +46,7 @@
See docs/format_internals.md for the formats acidcat walks.
"""

__version__ = "1.8.6"
__version__ = "1.8.7"

# dissection namespaces
from acidcat.core import probe # noqa: E402,F401
Expand Down
18 changes: 15 additions & 3 deletions src/acidcat/commands/audit.py
Original file line number Diff line number Diff line change
Expand Up @@ -278,6 +278,17 @@ def _run_one(args):
print(f"acidcat audit: {path}: {e}", file=sys.stderr)
return 2
size = os.path.getsize(path)
# a format acidcat recognises but only extracts from (a console ROM) is
# not "unknown", and `extract` is the verb that gets at its audio
extract_only = None
if not scanned:
from acidcat.core.infra import sniff
from acidcat.core.walk import _EXTRACT_ONLY
what = _EXTRACT_ONLY.get(sniff.sniff(path))
if what:
extract_only = what.split(" ", 1)[1]
label = label or extract_only
todo = "extract" if extract_only else "locate"

if chosen_format(args) == "json":
out = {
Expand Down Expand Up @@ -329,8 +340,9 @@ def _run_one(args):

if not scanned:
print(" FORENSICS not scanned -- no walker for this format")
print(" try: acidcat locate " +
os.path.basename(path) + " (finds embedded audio regardless)")
print(f" try: acidcat {todo} " + os.path.basename(path)
+ (" (recovers its samples)" if extract_only
else " (finds embedded audio regardless)"))
elif not other:
print(" FORENSICS nothing else flagged")
else:
Expand Down Expand Up @@ -378,6 +390,6 @@ def _run_one(args):
# "clean" is a claim about checks that ran. With no walker they did not,
# and the honest answer is that this verb had nothing to say -- `locate`
# still finds embedded containers in a format we cannot walk.
bits.append("clean" if scanned else "not analyzable -- no walker; try acidcat locate")
bits.append("clean" if scanned else f"not analyzable -- no walker; try acidcat {todo}")
print(f"\n VERDICT: {', '.join(bits) if bits else 'no structural fixes; review findings'}")
return _code(scanned, vios, findings, integ)
28 changes: 20 additions & 8 deletions src/acidcat/commands/inspect.py
Original file line number Diff line number Diff line change
Expand Up @@ -225,6 +225,12 @@ def _render_anomalies(findings, args):
print(f" {tag} {off} {f['rule']:16} {f['message']}")


def _at(offset):
"""a chunk's offset for the table; a node the walk does not place
(an SF2 preset, a tree node) has none"""
return "-" if offset is None else f"0x{offset:08x}"


def _render_table(filepath, fmt_label, chunks, file_warns, args, total=None,
shown_as=None):
file_size = os.path.getsize(filepath)
Expand All @@ -241,16 +247,17 @@ def _render_table(filepath, fmt_label, chunks, file_warns, args, total=None,
for i, c in enumerate(chunks):
idx = p("dim", f"[{c.get('_idx', i):>2}]")
cid = p("id", f"{c['id']:<5}")
off = p("dim", f"0x{c['offset']:08x}")
print(f" {idx} {cid} {off} {c['size']:<11,} {c['summary']}")
off = p("dim", _at(c["offset"]))
size = "-" if c["size"] is None else f"{c['size']:,}"
print(f" {idx} {cid} {off} {size:<11} {c['summary']}")

if not args.quiet:
for c in chunks:
if not c["fields"] and not c.get("rows"):
continue
print()
hdr_id = p("id", c["id"].strip())
hdr_meta = p("dim", f"@ 0x{c['offset']:08x} ({c['size']} bytes)")
hdr_meta = p("dim", f"@ {_at(c['offset'])} ({c['size'] if c['size'] is not None else '-'} bytes)")
print(f"{hdr_id} {hdr_meta}")
for fl in c["fields"]:
note = p("dim", f" {fl['note']}") if fl["note"] else ""
Expand All @@ -261,7 +268,8 @@ def _render_table(filepath, fmt_label, chunks, file_warns, args, total=None,
if o is not None else " ")
off_col = p("dim", off_col)
val = p("val", f"{fl['value']!s:<14}")
if args.show_hex and fl["off"] is not None:
if args.show_hex and fl["off"] is not None and (
c.get("payload_base") is not None or c["offset"] is not None):
# field offsets are measured from the chunk's payload base.
# RIFF/AIFF/RF64/MThd all have an 8-byte id+size header, so
# that is the default; formats with a different header (FLAC
Expand Down Expand Up @@ -329,18 +337,20 @@ def _full_chunk(chunk, filepath):
payload base, the raw region bytes as hex (capped), and every field's
absolute byte offset. `acidcat explore` needs nothing but this JSON."""
c = {k: v for k, v in chunk.items() if k != "_idx"}
pb = chunk.get("payload_base", chunk["offset"] + 8)
pb = chunk.get("payload_base")
if pb is None and chunk["offset"] is not None:
pb = chunk["offset"] + 8
c["payload_base"] = pb
fields = []
for f in chunk["fields"]:
f2 = _public_field(f)
# absolute file offset, so a field maps to raw[abs - offset]
f2["abs"] = pb + f["off"] if f["off"] is not None else None
f2["abs"] = pb + f["off"] if f["off"] is not None and pb is not None else None
fields.append(f2)
c["fields"] = fields
# only carry raw bytes for chunks that actually have positioned fields;
# audio-data regions are huge and have nothing to highlight.
if any(f["off"] is not None for f in chunk["fields"]):
if chunk["offset"] is not None and any(f["off"] is not None for f in chunk["fields"]):
n = min(chunk["size"], _FULL_RAW_CAP)
with open(filepath, "rb") as fh:
fh.seek(chunk["offset"])
Expand Down Expand Up @@ -649,7 +659,9 @@ def run(args):
# tuned on trackers broke silently on RIFF. --full has
# always emitted the absolute offsets; plain --json now
# does too, at no extra cost.
pb = c.get("payload_base", c["offset"] + 8)
pb = c.get("payload_base")
if pb is None and c["offset"] is not None:
pb = c["offset"] + 8
oc["payload_base"] = pb
fields = []
for f in c.get("fields", []):
Expand Down
6 changes: 5 additions & 1 deletion src/acidcat/commands/od.py
Original file line number Diff line number Diff line change
Expand Up @@ -333,7 +333,11 @@ def dim(t):
print(header)

for c in chunks:
base = c.get("payload_base", c["offset"] + 8)
if c.get("offset") is None:
continue # a tree node (an SF2 preset) has no bytes to dump
base = c.get("payload_base")
if base is None:
base = c["offset"] + 8
summary = c.get("summary", "")
title = _c("1;37", f"{str(c['id'])!r} @ 0x{c['offset']:08x} {c['size']:,} bytes", on)
print("\n" + title + (dim(" " + summary) if summary else ""))
Expand Down
12 changes: 10 additions & 2 deletions src/acidcat/core/forensics/anomalies.py
Original file line number Diff line number Diff line change
Expand Up @@ -486,10 +486,13 @@ def scan(filepath, fmt_label=None, chunks=None, warns=None):
# size-based check above cannot see it). Scans the last 64K+ from the end.
if not any(f["rule"] == "polyglot" for f in findings):
with open(filepath, "rb") as f:
is_zip = f.read(4) == b"PK\x03\x04"
f.seek(max(0, size - 66000))
tail = f.read()
idx = tail.rfind(b"PK\x05\x06")
if idx >= 0 and len(tail) - idx >= 22:
# a file that IS a zip (.xpn, .labx, .multisample) ends with its own
# end record; that is the archive's index, not one appended to it
if idx >= 0 and len(tail) - idx >= 22 and not is_zip:
findings.append({"severity": "alert", "offset": (size - len(tail)) + idx,
"rule": "polyglot",
"message": "possible polyglot: appended ZIP archive "
Expand Down Expand Up @@ -748,8 +751,13 @@ def scan(filepath, fmt_label=None, chunks=None, warns=None):
# branched on the label, and it made the guard depend on the label being
# non-empty: a walker that returned "" got a spurious embedded-Ogg
# finding on every ordinary Ogg file.
# an archive's members are whole files by design (a .multisample zone
# is a WAV in a zip), and an SF3 stores every sample as an Ogg stream
wanted = [c for c in _EMBEDDED_MEDIA
if not (c[0] == b"OggS" and fmt_id in _OGG_IDS)]
if not (c[0] == b"OggS" and fmt_id in _OGG_IDS)
and not (c[0] == b"OggS" and fmt_id == "sf2")]
if own == b"PK\x03\x04":
wanted = []
hits = _find_embedded(filepath, wanted, own) if wanted else {}
for magic, _form_at, _form, name in wanted:
at = hits.get(magic)
Expand Down
22 changes: 22 additions & 0 deletions src/acidcat/core/forensics/provenance.py
Original file line number Diff line number Diff line change
Expand Up @@ -210,6 +210,25 @@ def _ffmpeg_rf64_junk(chunks):
return None


def _ableton_junk_quote(chunks):
"""Ableton Live writes a 28-character quote ("The sleeper must awaken",
"Why r u using a hex editor?") into the RF64 reservation of the WAVs it
renders; measured in two unrelated Live projects. A later tool may write
the RIFF size over its first 8 bytes, and the tail still names Live. The
walker has decoded the text; this reads its summary."""
first = chunks[0] if chunks else {}
if str(first.get("id", "")).strip() != "JUNK":
return None
s = first.get("summary", "")
if s.startswith("RF64 reservation (ds64) holding text"):
basis = "a quote in the 28-byte RF64 reservation"
elif "written over a text" in s:
basis = "the tail of a quote under a later tool's RF64 sizes"
else:
return None
return {"tool": "Ableton Live", "confidence": "likely", "basis": basis}


def _structural(label, chunks, data):
out = []
if "MP3" in label or "MPEG" in label:
Expand All @@ -232,6 +251,9 @@ def _structural(label, chunks, data):
ff = _ffmpeg_rf64_junk(chunks)
if ff:
out.append(ff)
live = _ableton_junk_quote(chunks)
if live:
out.append(live)
return out


Expand Down
19 changes: 8 additions & 11 deletions src/acidcat/core/formats/mdx.py
Original file line number Diff line number Diff line change
Expand Up @@ -254,23 +254,20 @@ def looks_like_mdx(raw, filesize=None):

Everything before the offset table is variable-length text, so the only
thing that can identify an MDX is whether the arithmetic lands: a title
terminator, a NUL, and a table that resolves to a legal channel count with
every offset inside the file.

`filesize` matters. The offsets routinely point past the few kilobytes a
sniffer reads, so checking them against len(raw) rejects almost every real
tune when handed a truncated head -- which is exactly what happened, and
the bound has to be the FILE's length rather than the buffer's.
terminator, a NUL, and a table that resolves to a legal channel count.
Offsets that run past the end mark a cut file, not a different format;
`filesize` is kept for callers and no longer consulted.
"""
h = parse_header(raw)
if not h["ok"]:
# a packed module is an MDX whose body cannot be walked, not a file
# that is something else
return bool(h.get("packer"))
n = len(raw) if filesize is None else filesize
if not 0 <= h["voice_abs"] <= n:
return False
return all(0 <= a <= n for a in h["mml_abs"])
# Offsets past the end of the file do not make it something else: a cut
# rip keeps a whole header (terminator, bank name, a 9- or 16-channel
# table) and the walker reports each offset that dangles. 16 real modules
# are cut like that, and no other file in the corpus has such a header.
return True


def voice_count(raw, h):
Expand Down
Loading
Loading