Skip to content

About

Two custom fuzzers for the GZDoom engine with crash triage: 15,448 logged runs, 1,433 crashing WAD inputs, and an AFL++ control arm that found none.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

GZDoom Fuzzing Framework

Two custom Python fuzzers, a corpus builder, and a crash minimizer, pointed at the GZDoom game engine. Across 15,448 logged runs over roughly 28 hours, the corpus PWAD fuzzer produced 1,433 crashing inputs (20.6% of its runs: 1,074 SIGABRT under AddressSanitizer, 359 timeouts), while the PK3 script fuzzer drove GZDoom to a non-zero exit on 8,476 of 8,477 runs but deduplicated down to only 2 distinct crash signatures. A stock AFL++ comparison arm run against the same engine for 10 hours 20 minutes found zero crashes at 0.47% bitmap coverage.

That last contrast is the actual result: the same target, the same time budget, and the naive corpus-mutation approach reached crash states that a coverage-guided industry-standard fuzzer never reached, because AFL++'s bit-level mutations destroyed the WAD structure before the engine would parse the file at all.

Built for CSE 5472 (Information Security Projects) at Ohio State, fall 2025.

Team and my role

This was a three-person project with Sam Halsey and Rushi Bhatt. We worked through every component together rather than splitting it into separate ownership, so I contributed across all of it: the extractors, both fuzzers, the triage tooling, and the campaign analysis. The writing in this README is mine; the findings it reports come from our team's evaluation writeup and the metrics our runs produced, and I have marked which is which throughout.

Results

Every number below is computed from the two CSVs in data/, which are the raw per-iteration logs from the actual campaigns. Nothing here is estimated.

Corpus PWAD fuzzer

Ran 2025-11-28 18:24 to 2025-11-29 08:57.

Metric Value
Logged iterations 6,971
Wall clock 14 h 33 m
Throughput ~479 execs/hour
Clean exits (rc 0) 5,538 (79.4%)
SIGABRT (rc -6) 1,074 (15.4% of runs, 74.9% of crashes)
Timeouts at the 10 s limit 359 (5.2% of runs, 25.1% of crashes)
Crashing inputs total 1,433 (20.6%)
First crash iteration 1, at 2.94 seconds
Mean duration, clean run 8.02 s
Mean duration, aborting run 4.20 s

The crash rate never decayed. The first 1,000 iterations produced 201 crashing inputs and the last 1,000 produced 216, which is why our evaluation report described discovery as linear rather than saturating.

PK3 ZScript/DECORATE fuzzer

Ran 2025-11-25 20:00 to 2025-11-26 09:30.

Metric Value
Logged iterations 8,477
Wall clock 13 h 30 m
Throughput ~628 execs/hour
Exit code 255 6,530 (77.0%)
SIGABRT (rc -6) 1,945 (22.9%)
Clean exits 1
Timeouts 1
Distinct triage keys reached 2 (at iterations 5 and 1,099)

This campaign is a lesson in oracle design more than a bug haul. The harness flags a crash on any non-zero return code, and GZDoom exits 255 on an ordinary fatal script error, so three quarters of these "crashes" are the engine correctly rejecting malformed ZScript. The 1,945 SIGABRTs are the interesting subset, and the triage counter collapsed all of them into 2 signatures.

AFL++ comparison arm

Our team also ran AFL++ 4.09c against an AFL-instrumented GZDoom build as a control. This was a single comparison run, not a tuned campaign, and the numbers below are read off the AFL++ status screen captured in our final presentation rather than from a logged dataset.

Metric Value
Run time 10 h 20 m 34 s
Total executions 18.2k (0.50 execs/sec)
Saved crashes 0
Saved hangs 32
Total timeouts 366
Map density 0.47%
Stability 99.22%
Cycles completed 0

What was actually built

src/extract_map_pwad.py    single map lump group out of an IWAD into a standalone PWAD
src/batch_extract_maps.py  same, swept across every MAP##/E#M# marker in a WAD
src/simple_fuzzer.py       first pass: byte mutations on one seed WAD
src/corpus_fuzzer.py       corpus-driven WAD fuzzer with metrics logging and crash triage
src/pk3_text_fuzzer.py     unzip PK3, mutate zscript.txt/DECORATE, repackage, run
src/reduce_wad.py          ddmin-style minimizer, zeroes map-lump regions while the crash holds

The extractors exist because of a practical blocker our report records: a full IWAD carries hundreds of lumps, so feeding one to a fuzzer wastes nearly every mutation on data the map loader never reaches. Pulling each map out as its own small PWAD is what made a useful seed corpus possible.

The two fuzzers share a harness shape. Mutate a seed, write it to a temp directory, launch GZDoom headless with a stable Freedoom Phase 2 IWAD and +map MAP01 +quit, classify the result, and append one row per iteration to a CSV. The corpus fuzzer runs against a Clang build of GZDoom instrumented with AddressSanitizer and UndefinedBehaviorSanitizer under abort_on_error=1, which is what converts a silent memory error into the SIGABRT the harness can see.

The one crash frame in our materials

The only individual crash log preserved in our deliverables is a screenshot in the final presentation, showing RETURN_CODE: -6 on a run that died at:

src/common/filesystem/source/files_decompress.cpp:847
FileSys::DecompressorLZSS::Read(...)

That is a failure inside the LZSS decompressor in GZDoom's filesystem layer, reached while loading a mutated archive. I am reporting it as what it is: one observed abort site, captured in a slide, not a diagnosed and reported vulnerability.

What I would do differently

Fix the triage key. This is the biggest flaw in the project and it inflates the headline number. corpus_fuzzer.py computes its crash key like this:

def triage_key(mutated_bytes: bytes, log_text: str) -> str:
    h = hashlib.sha256()
    h.update(mutated_bytes)          # <- every mutated file is unique, so every key is unique
    for line in log_text.splitlines():
        if "AddressSanitizer" in line or "SUMMARY:" in line or "FScanner::ScanString" in line:
            h.update(line.encode("utf-8", "ignore"))
    return h.hexdigest()

Because the mutated input bytes go into the hash, two runs that abort at the identical source line with the identical stack still produce different keys. "1,433 unique crashes" is therefore 1,433 crashing inputs, not 1,433 bugs, and the true bug count is unknown and certainly far smaller. Our evaluation report caught the symptom, noting the linear growth curve implied the deduplication was "overly sensitive," and attributed it to hashing dynamic elements. Reading the code back a year later, the specific cause is the h.update(mutated_bytes) line. The fix is to hash only a normalized stack signature: top N frames, with addresses and file paths stripped.

The PK3 fuzzer is the control that proves the point. It used a key derived from output alone and collapsed 8,476 abnormal exits into 2 signatures. Same campaign scale, different key, three orders of magnitude different answer.

Separate the oracle from the exit code. Treating any non-zero exit as a crash is what buried the PK3 results under 6,530 script-error exits. A real oracle would classify on the sanitizer report, not on returncode != 0.

Actually run the minimizer. reduce_wad.py was written and is included here, but our deliverables contain no minimized reproducer, so I cannot claim a before/after size. Minimizing even a handful of the 1,433 crashing WADs would have turned a pile of inputs into a small set of understandable bugs.

Triage before scaling up. Our report is candid that the triage infrastructure landed last, and that until it existed we were mostly confirming the components ran. Building the classifier before spending 28 hours of compute would have made those hours worth more.

Limitations

Stated plainly, because these bound what the numbers above mean.

  • No crash was diagnosed to a root cause. No ASan report, stack trace, or minimized reproducer is preserved in our materials beyond the single slide frame quoted above. The crash counts are counts of crashing inputs.
  • No vulnerability was disclosed, and none should be inferred. Nothing here was confirmed as exploitable, tested against a current GZDoom release, or reported upstream. The engine was run headless on synthetic maps in a throwaway WSL2 environment.
  • The two campaigns are not directly comparable. They ran on different builds, different seeds, and different targets, so the 20.6% versus 99.99% figures measure different things.
  • The AFL++ arm was a single untuned run. 0 crashes over one 10-hour run at 0.50 execs/sec is evidence about that configuration, not a general claim that AFL++ cannot fuzz GZDoom. A dictionary, a proper -x token file, or a persistent-mode harness would very likely change the result.
  • corpus_fuzzer.py and the CSV disagree on schema. The committed CSV header lists 9 columns while its rows carry 11. See data/README.md; the file is published exactly as the run produced it.
  • Earlier drafts of our own writeup overstate the findings. An intermediate deliverables draft listed specific bug classes (use-after-free in class inheritance resolution, heap overflow in ZScript lexical scanning, and others) and put the PK3 crash rate at 50%. The metrics logs do not support those statements and our final report dropped them, so they are not repeated here.

Running it

Requires Linux (we used Ubuntu on WSL2), Python 3, an IWAD, and a GZDoom build. For the sanitizer oracle to work, GZDoom must be compiled with Clang and -fsanitize=address,undefined; the flatpak build referenced as a default in some scripts will run inputs but cannot report memory errors.

Build a seed corpus from an IWAD:

python3 src/batch_extract_maps.py --iwad /path/to/freedoom2.wad --outdir corpus/

Run the corpus fuzzer against a sanitizer build:

python3 src/corpus_fuzzer.py \
  --corpus corpus/ \
  --outdir corpus_output/ \
  --iterations 1000 \
  --timeout 10 \
  --stable-iwad /path/to/freedoom2.wad \
  --gzdoom-bin /path/to/gzdoom/build/gzdoom

Run the PK3 script fuzzer against a seed archive containing zscript.txt or DECORATE:

python3 src/pk3_text_fuzzer.py --seed seed.pk3 --iterations 500 --iwad /path/to/freedoom2.wad

The environment the campaigns used:

ASAN_OPTIONS=abort_on_error=1:detect_leaks=0:alloc_dealloc_mismatch=1:symbolize=1
UBSAN_OPTIONS=print_stacktrace=1
SDL_AUDIODRIVER=dummy
SDL_VIDEODRIVER=wayland, falling back to x11

seeds/seed-map01.wad is MAP01 extracted from Freedoom Phase 2 by extract_map_pwad.py, included so the tools can be run without sourcing an IWAD first.

Repository contents

src/          the six scripts, as written for the project
data/         raw per-iteration metrics from both campaigns, plus a schema note
seeds/        MAP01 extracted from Freedoom Phase 2
docs/         evaluation summary drawn from the team's final report

Attribution and licensing

The code and this README are licensed MIT. seeds/seed-map01.wad is derived from Freedoom Phase 2 v0.13.0, distributed under the Freedoom BSD-style license, and remains under that license. GZDoom itself is not included here; it is a separate project under GPL.

Course context is given generically. No assignment text, instructor-provided code, or course dataset appears in this repository.

About

Two custom fuzzers for the GZDoom engine with crash triage: 15,448 logged runs, 1,433 crashing WAD inputs, and an AFL++ control arm that found none.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages