llm_model_inspector.py reads a local model container and produces a detailed Microsoft Word (.docx) report without running inference. Its native GGUF and SafeTensors readers inspect only bounded headers, metadata, tensor directories, shapes, types, and offsets; they do not materialize tensor payloads.
Warning
This tool reports container structure and inferred architecture; it does not verify model behavior or prove that a model is safe. Never use the optional PyTorch mode on an untrusted checkpoint.
The report includes:
- model/container identity and file-region maps;
- declared and inferred architecture, with confidence and evidence;
- layers, embedding/FFN widths, attention and KV heads, head dimension, context, vocabulary, experts, normalization, QKV layout, and position encoding when discoverable;
- diagrams of the end-to-end pipeline, repeated block, file anatomy, and parameter distribution;
- parameter and storage totals by subsystem, type/quantization, rank, and layer;
- configuration, tokenizer, and container metadata;
- a landscape tensor appendix with on-disk and framework-style shapes, role, layer, dtype, storage, offset, and shard;
- explicit warnings for partial models, conflicts, unknown encodings, or structural inconsistencies.
| Input | Inspection behavior |
|---|---|
| GGUF v2/v3 | Native metadata-only parser; little/big endian; current GGML types; split files auto-discovered from -00001-of-NNNNN names |
.safetensors |
Native header-only parser with strict shape, range, overlap, and byte-size validation |
*.safetensors.index.json |
Resolves and cross-checks every local shard; adjacent indexes are auto-discovered when a shard is selected |
| Hugging Face model directory | Discovers one unambiguous SafeTensors index/file or GGUF model and reads adjacent config.json / tokenizer_config.json |
*.bin.index.json |
Safe manifest-only report by default: tensor names and shard membership, but no pickle payloads |
.bin, .pt, .pth, .ckpt |
Rejected by default; trusted-file mode requires PyTorch ≥2.10 and uses weights_only=True with PyTorch's documented FakeTensorMode in a timeout/resource-limited subprocess |
Old GGML files, ONNX, HDF5, and arbitrary pickle formats are intentionally not guessed. Converting to GGUF or SafeTensors gives the fullest and safest report.
Python 3.10 or newer is required.
Install from a source checkout:
git clone https://github.com/deanhorak/llm_model_inspector.git
cd llm_model_inspector
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install .
llm-model-inspector --versionOn Windows, activate the environment with .venv\Scripts\activate.
The standalone-script workflow is also supported:
python -m pip install -r requirements.txt
python llm_model_inspector.py --versionThe standard report needs only python-docx and Pillow. PyTorch is not a normal dependency; install a patched PyTorch 2.10 or newer separately only if you explicitly need --pytorch-mode trusted-weights-only for a checkpoint you trust.
llm-model-inspector /models/model.gguf
llm-model-inspector /models/model.safetensors -o architecture.docx
llm-model-inspector /models/huggingface-directory --paper a4
llm-model-inspector /models/pytorch_model.bin.index.jsonEquivalent invocations are python -m llm_model_inspector ... and python llm_model_inspector.py ....
The default output is <input>_structure_report.docx in the current directory. Existing files are never replaced unless --overwrite is supplied.
Useful options:
--no-diagrams Omit PNG diagrams; Pillow becomes optional
--max-tensors-in-report N Limit appendix rows (default 10,000); 0 requests all
--max-metadata-rows N Limit each metadata table (default 2,000); 0 requests all
--include-paths Include absolute source paths (off by default)
--max-header-mib N Raise/lower the bounded header limit
--paper letter|a4 Select page size
--pytorch-mode trusted-weights-only
Inspect a trusted checkpoint with patched PyTorch
--torch-timeout SECONDS Bound each PyTorch helper process
--torch-memory-gib GIB POSIX address-space cap for that helper
The default 10,000-row ceiling keeps pathological or extremely large MoE inventories from producing an unusable Word file while preserving complete architecture and aggregate statistics. Use --max-tensors-in-report 0 only when you explicitly want every tensor row.
See the example report, generated from a tiny synthetic Llama-like fixture that contains no third-party model weights or data.
- GGUF/SafeTensors payload bytes are not loaded.
- JSON uses duplicate-key rejection, finite-number parsing, depth/item limits, and bounded file sizes.
- Shard paths must be relative, local, present, and resolve inside the index directory; traversal, URI, absolute-path, and symlink escapes are rejected.
- Model repositories are not imported, network access is not used, and
trust_remote_codeis never enabled. - Plain
torch.load()andweights_only=Falseare never used. The optional mode hard-requires PyTorch ≥2.10 because olderweights_onlyloaders have published code-execution vulnerabilities. PyTorch still says never to load untrusted data, so this mode is only for a checkpoint you trust and is not a security sandbox. See GHSA-63cw-57p8-fm3p and thetorch.loadwarning. - Tensor/config strings are sanitized before insertion into Office XML.
Reports omit absolute source paths by default. They can still contain model-controlled metadata, tensor names, authorship, license, and URL fields. Review a report before sharing it. --include-paths intentionally adds absolute local paths.
Configuration values are reported as declared. GGUF metadata fields are reported as exact. Calculations are derived, while tensor-name/shape conclusions are marked strong or heuristic. The report keeps conflicting evidence visible.
GGUF dimensions are stored fastest-dimension-first, so the inventory shows both their raw on-disk order and the reversed framework-style order. SafeTensors and PyTorch shapes already use framework row-major order. Parameter totals describe stored scalar values; they do not prove runtime weight sharing or whether an absent output tensor is tied/synthesized.
- Architecture conclusions can be heuristic when configuration or canonical tensor names are absent.
- The tool does not validate numerical weights, model quality, or executable model behavior.
- Native parsing reads bounded metadata/header regions, although split models require opening every shard.
- A complete tensor appendix can make very large DOCX files; the default report samples at most 10,000 rows.
- Resource limits reduce denial-of-service risk but are not a hostile-file sandbox.
| Code | Meaning |
|---|---|
0 |
Report generated successfully |
2 |
Usage, unsupported-format, malformed-input, safety, or output error |
130 |
Interrupted by the user |
python -m pip install -e ".[dev]"
python -m unittest discover -s tests -v
python -m compileall -q llm_model_inspector.py tests
ruff check .
python -m buildThe CLI reference, implementation architecture, report contents, and detailed format notes are in docs/. Before contributing, read CONTRIBUTING.md, SECURITY.md, and CODE_OF_CONDUCT.md.
No open-source license has been granted for this repository. Copyright rights are reserved by default; contact the maintainer before redistributing or incorporating the code elsewhere.