Drop pdf files from the "MCRO Evidentiary Dataset" into the folder and all internal objects are extracted, hashed, and inserted into a csv table along with digital signature data (pdfsig) and metadata author and creator (exiftool)
The "MCRO Evidentiary Dataset" consists of 3,601 total PDF files which are purported to be authentic Minnesota Judicial Records - BUT THEY ARE NOT. Machine generated, mass forgery spanning 163 supposed cases in just this set.
Guertin v. Walz, et al. 25-cv-2670-PAM-DLM, D. Minn 2025: https://www.MnCourtFraud.com/docket/70633540/guertin-v-walz/ https://www.courtlistener.com/docket/70633540/guertin-v-walz/
Matthew Guertin v. Tim Walz, et al. 25-2476, 8th Cir. 2025: https://www.courtlistener.com/docket/70929923/matthew-guertin-v-tim-walz/
Among the fraud is:
40 hash matched USPS "Returned Mail" filings: https://MnCourtFraud.com/recap/gov.uscourts.mnd.226147/gov.uscourts.mnd.226147.59.0.pdf https://storage.courtlistener.com/recap/gov.uscourts.mnd.226147/gov.uscourts.mnd.226147.59.0.pdf
1,165 cloned judicial signatures (hash matched): https://MnCourtFraud.com/recap/gov.uscourts.mnd.226147/gov.uscourts.mnd.226147.57.0.pdf https://storage.courtlistener.com/recap/gov.uscourts.mnd.226147/gov.uscourts.mnd.226147.57.0.pdf
371 cloned, judicial timestamps (hahs matched: https://MnCourtFraud.com/recap/gov.uscourts.mnd.226147/gov.uscourts.mnd.226147.58.0.pdf https://storage.courtlistener.com/recap/gov.uscourts.mnd.226147/gov.uscourts.mnd.226147.58.0.pdf
27 Cloned "Correspondence for Judicial Approval" filings: https://MnCourtFraud.com/recap/gov.uscourts.mnd.226147/gov.uscourts.mnd.226147.19.0.pdf https://storage.courtlistener.com/recap/gov.uscourts.mnd.226147/gov.uscourts.mnd.226147.19.0.pdf
Machine generated batches of fraudulent "Finding of Incompetency and Order" filings with clone X.509 signatures Same signing time, and same byte ranges, with 12 clones in a single group spanning different case numbers: https://MnCourtFraud.com/recap/gov.uscourts.mnd.226147/gov.uscourts.mnd.226147.15.2.pdf
The "MCRO Evidentiary Dataset" of 3,601 AI gneerated Minnesota court records (all with valid, full doc, X.509 court sigs) can be downloaded at: https://MnCourtFraud.com/File/2017.zip - /2023.zip
Backup mirrors available as file embedded, digitally signed PDF wrappers: https://Matt1Up.Substack.com/p/evidence https://MnCourtFraud.substack.com/p/mcro-files
This repository/script watches a pdf/ folder for new PDF files and, for each PDF:
- Explodes it into per-object files (images, fonts, etc.) using
mutool extract. - Hashes every extracted object (SHA‑256).
- Appends a row per object to
objects.tsvwith rich metadata (see schema). - Copies each unique object (by hash) into
hashed-objects/named<sha256>.<ext>(deduplicated by content). - Maintains a
processed.tsvledger so each PDF is processed once. - Maintains a
hash-count.tsvsummary of unique object hashes and counts. - Parses MCRO-style filenames to extract
Case Number,Filing Type,Filing Date. - Extracts PDF metadata (Author/Creator via
exiftool) and digital-signature fields (viapdfsig) — signed ranges, signer CNs, and signing times (normalized).
The script is safe with spaces in filenames, uses locks to avoid races, and supports both one-shot catch‑up and monitor mode.
# 1) Place the script in a working directory and make it executable
chmod +x pdf_object_hasher.sh
# 2) Ensure directory layout exists (the script will also create them if missing)
mkdir -p pdf pdf-objects hashed-objects
# 3) Run once (catch-up any new PDFs in ./pdf/)
./pdf_object_hasher.sh
# 4) Or run continuously (watch for new PDFs)
./pdf_object_hasher.sh --monitorWhen running in monitor mode, if
inotifywaitis not installed, the script will poll every 5 seconds.
<working-dir>/
├─ pdf/ # Drop PDFs here (script reads from this folder)
├─ pdf-objects/ # Per-PDF extracted objects (1 subfolder per PDF; safe name)
│ └─ <safe>/
│ ├─ image-0001.png
│ ├─ font-0003.ttf
│ └─ .processed.sha # stamp equal to the PDF’s SHA256 (used to prevent reprocessing)
├─ hashed-objects/ # Unique blobs copied by content hash → <sha256>.<ext>
├─ objects.tsv # Main table (one row per extracted object, schema below)
├─ hash-count.tsv # <sha256>\t<count> summary (rebuilt after each PDF)
├─ processed.tsv # <sha256(pdf)>\t<filename>\t<bytes>\t<mtime_epoch>\t<processed_utc_iso>
└─ .locks/
└─ inflight/ # in-flight guards: one file per PDF SHA during processing
Header (in order):
- Case Number — from
MCRO_filename (blank if not an MCRO file) - Filing Type — from
MCRO_filename (blank if not an MCRO file) - Filing Date — from
MCRO_filename (blank if not an MCRO file) - SHA256 Hash Value — object content hash (used for deduplication)
- Pdf File Name — original PDF filename
- Pdf Internal Object Path — path relative to
pdf-objects/ - Object Type — lowercase file extension (with leading dot), if present (e.g.,
.png,.ttf) - Font Name — best‑effort (from
otfinfoorfc-scan), blank if unknown - Sig #1 Common Name — signer CN from
pdfsig(signature block #1) - Sig #2 Common Name — signer CN from
pdfsig(signature block #2) - Author — from
exiftool -Author - Creator — from
exiftool -Creator - Sig #3 Common Name — signer CN from
pdfsig(signature block #3) - Sig #4 Common Name — signer CN from
pdfsig(signature block #4) - Sig #1 Signing Time — normalized
YYYY-MM-DD HH:MM:SS - Sig #2 Signing Time — normalized
YYYY-MM-DD HH:MM:SS - Sig #3 Signing Time — normalized
YYYY-MM-DD HH:MM:SS - Sig #4 Signing Time — normalized
YYYY-MM-DD HH:MM:SS - Sig #1 Byte Ranges — e.g.,
[0 - 157337], [169389 - 200740] - Sig #2 Byte Ranges — as above
- Sig #3 Byte Ranges — as above
- Sig #4 Byte Ranges — as above
For non‑MCRO filenames, columns 1–3 remain blank.
If a PDF lacks signatures or metadata, the relevant columns are left blank.
A PDF’s signature/author/creator values repeat per-object-row (expected).
Example header row:
Case Number Filing Type Filing Date SHA256 Hash Value Pdf File Name Pdf Internal Object Path Object Type Font Name Sig #1 Common Name Sig #2 Common Name Author Creator Sig #3 Common Name Sig #4 Common Name Sig #1 Signing Time Sig #2 Signing Time Sig #3 Signing Time Sig #4 Signing Time Sig #1 Byte Ranges Sig #2 Byte Ranges Sig #3 Byte Ranges Sig #4 Byte Ranges
- Built from column 4 (SHA256) of
objects.tsv. - Format:
<sha256>\t<count>sorted by descendingcount.
<sha256(pdf)>\t<pdf_filename>\t<bytes>\t<mtime_epoch>\t<processed_utc_iso>
The script also writes a stamp file
pdf-objects/<safe>/.processed.shacontaining the same SHA to reinforce one‑time processing.
If a file starts with MCRO_, the script splits the filename by underscores:
MCRO_<Case Number>_<Filing Type>_<Filing Date>_...
Case Number= string after the first underscore and before the second.Filing Type= string after the second underscore and before the third.Filing Date= string after the third underscore and before the fourth.
If the filename does not start with MCRO_, these three columns are left blank.
pdfsig(Poppler) is used to collect up to four signature blocks:- Signer Common Name (CN) →
Sig #X Common Name - Signing Time → normalized to
YYYY-MM-DD HH:MM:SS - Signed Ranges → captured verbatim per block
- Signer Common Name (CN) →
exiftoolprovidesAuthorandCreator.
If a given field isn’t present in the PDF, the corresponding column stays blank.
- One-time processing — A PDF is considered processed if its SHA256 exists in
processed.tsv, or the per‑PDF stamp file matches. - In-flight lock — While a PDF is being processed, an inflight lock prevents re‑entry.
- Space-safe — Uses null‑delimited
findand robust quoting. - Quiet-file wait — Small delay to ensure a PDF is fully written before processing.
- Dedup store — Any object (by hash) is copied once into
hashed-objects/. If the same hash reappears with a different extension, the first-seen copy is kept.
Required:
mutool(frommupdf-tools) — object extractionsha256sum(GNU coreutils) — hashingbash,awk,sed,find,stat,date— standard UNIX tools
Metadata / signatures:
exiftool— Author/Creatorpdfsig(frompoppler-utils) — signature details
Optional (for font names in objects.tsv):
otfinfo(fromlcdf-typetools) orfc-scan(fromfontconfig)
Optional (live watch, otherwise polling is used):
inotifywait(frominotify-tools)
sudo apt-get update
sudo apt-get install -y mupdf-tools poppler-utils exiftool inotify-tools fontconfig lcdf-typetools coreutilsbrew install mupdf-tools poppler exiftool coreutils fontconfig lcdf-typetools
# Optional: fswatch for watching if you prefer; the script will poll if inotifywait is absentmacOS note:
date -dis GNUdate. If needed, install GNU coreutils and ensuregdateis available; the script usesdate -d, so consider aliasingdate=gdateor adjusting your PATH.
sudo pacman -S --needed mupdf poppler exiftool inotify-tools fontconfig lcdf-typetools coreutils# Help
./pdf_object_hasher.sh --help
# One-time processing of anything new in ./pdf/
./pdf_object_hasher.sh
# Continuous watch (reacts to each new PDF)
./pdf_object_hasher.sh --monitorIf you intentionally need to reprocess a PDF (e.g., after edits):
- Remove its line from
processed.tsv(match by SHA in column 1). - Remove the stamp:
rm pdf-objects/<safe>/.processed.sha - (Optional) Remove its
pdf-objects/<safe>/folder if you want a clean re‑extract. - (Optional) Leave
hashed-objects/as-is to preserve dedupe, or remove specific hashes if appropriate.
After changes, re-run the script (catch-up or monitor mode).
The script automatically rebuilds counts after each PDF, but you can regenerate manually:
# Column 4 is the SHA256 in the new schema
tail -n +2 objects.tsv | awk -F'\t' 'NF>=4{print $4}' | sort | uniq -c | sort -nr \
| awk '{print $2 "\t" $1}' > hash-count.tsv- “Polling every 5s”: Install
inotifywaitif you want immediate reactions. - No signature columns populated: The PDF has no (or fewer) digital signatures, or
pdfsigisn’t installed. - Author/Creator blank: The PDF lacks those metadata fields, or
exiftoolisn’t installed. - Time parsing quirks: On non-GNU
date, times may not normalize; install GNUdate(coreutils) and ensure it’s used. - Spaces/weird paths: Fully supported. If you see issues, confirm you are running Bash.
- Idempotency:
processed.tsv+ per-PDF stamp file + inflight lock ensure one-time processing per unique PDF content. - Atomic-ish writes: temp files are moved into place to reduce risk of partial writes.
- Extensible: You can add more columns/sources; new columns will simply append to
objects.tsvrows.
Use freely for forensic analysis and research. No warranty.