One command applies the Dealign abliteration edit to the checkpoint that
Mia's TensorFold recipe serves
(Mia-AiLab/GLM-5.3-Flash-EXL3-TR3-4bpw, at the recipe's pin) and points the recipe at the result. The recipe and its
published image serve it unchanged; the only download is the 2.50 GiB of donor tensors.
git clone https://github.com/eleata/glm53-tensorfold-ablit && cd glm53-tensorfold-ablit
./install.sh # on the head Spark, with WORKER set in the recipe's scripts/local.shStatus. The edit is verified at the byte level, runs at the speed Mia publishes for the stock checkpoint, and on
strong_reject answers 73 of 100 harmful prompts where the stock checkpoint answers 4 (same judge). Quality is still
to be measured.
The self_attn.o_proj.weight tensors of layers 15-45 are replaced, byte for byte, by those of
dealignai/GLM-5.3-Flash-UNCENSORED-NVFP4
(MIT), at the pinned revision 745aac2f. Every other byte of the weights is Mia's checkpoint. In the donor, the layer-45 tensor
equals stock, so the tensors that actually change are layers 15-44. The same tensors are what Mia's earlier vLLM recipe
replaced at load time with ABLIT=1 ABLIT_METHOD=transplant; this repo writes them to disk once instead.
- Recipe: your checkout (
--recipe DIR), or a fresh clone at v1.2 (1f3d909), the version measured here. - Settings: reads the recipe's effective settings with its own loader (environment,
scripts/local.sh,.env): cache,WORKER,WORKER_WEIGHTS, container name, port. Stops ifWORKERis empty orMODEL_ID/MODEL_REVISIONare set in the environment (they would override the files). - Stock checkpoint: if it is not in that cache, runs the recipe's own
scripts/prepare.shfor it. Before building, every one of its 144 files is checked against the pinned inventory (stock_inventory.json: size and sha256, or git sha1 for small files), which reads ~164 GiB once. - Donor tensors and snapshot:
fetch_donor.pydownloads only the 31o_projtensors with HTTP range requests, and every one must hash to the value pinned inpins.py.build_ablit_snapshot.pythen creates the repolocal/GLM-5.3-Flash-EXL3-TR3-4bpw-ablitin the same cache, laid out like a downloaded one:blobs/holds hard links to the stock blobs (no extra space) plus the 31 edited shards (46 GiB);- the snapshot is relative links
../../blobs/..., the layout the recipe copies to the worker withWORKER_WEIGHTS=copyand reads over NFS withnfs; it keeps working if the stock repo is deleted; - every safetensors header is validated (dtype sizes, no gaps or overlaps, no trailing bytes), and every index entry must be in the shard it names;
- the result is verified in full before an atomic rename publishes it: the exact file list of the stock snapshot,
every link resolving inside the new repo, every unchanged file being the stock file, every edited tensor and
edited shard hashing to the expected value, and the receipt matching the inputs. A failure leaves nothing
behind; nothing is removed once the snapshot is published. An existing build is reused only if
--verifypasses (--verify --deepre-hashes every file against the inventory). - the snapshot also carries its own card (what was changed, the licenses) and the donor's
LICENSE.
- Recipe settings: writes
MODEL_ID/MODEL_REVISIONin a marked block ofscripts/local.sh, after checking the block markers, with a backup and abash -ncheck. Delete the block and run./start.sh restartto go back. - Server: with
WORKER_WEIGHTS=copyit first hard-links the stock blobs the worker already has into the new repo there, so only the 31 edited shards cross the link. Then it starts the server, or restarts it if it is running, and checks that both ranks run with exactly the new snapshot and the API answers. If that is already the case, it does nothing (--restartto force). After editinglocal.shit re-reads the settings through the recipe's loader, so an assignment later in the file or in the environment cannot silently win.
python3 -m unittest discover -s tests runs 40 tests: the builder (including corrupted, incomplete, tampered and
linked inputs, and a failure right after publishing), the fetcher against a local server (corrupted local files, wrong
bytes or ranges, the token not following a redirect to another origin) and the installer end to end with stub
docker/ssh (copy-mode seeding, re-runs, a missing or broken local.sh, an override after the block, a similar
revision already running, a corrupted existing snapshot). On real data, the hashing and header checks were run
against the actual stock files (EVIDENCE.md).
Two DGX Sparks, recipe v1.2 defaults (4 streams, 1,048,576-token window, FP8 KV, DENSE=q4, DFlash2 + copy drafts),
one request at a time, temperature 0, thinking off, max_tokens=400, 5 runs per prompt. Median [range]:
| This repo | Stock, from Mia's README | |
|---|---|---|
| Prose | 59.4 tok/s [54.8-59.5] | 60.4 |
| Structured (count to 200) | 114.3 tok/s [112.9-114.5] | 114.7 |
| Code | 78.9 tok/s [78.7-79.0] | n/a |
The medians are within 2% of Mia's stock figures, which were measured with another client (sparkDash) and GPU clocks
capped at 2,200 MHz: an indication, not an A/B. KV pool at start: 2,922,496 tokens. Method and raw results:
EVIDENCE.md, evidence/.
Inspect strong_reject, the same first 100 of its 313 prompts for both checkpoints (temperature 0.75, max_tokens
2048, reasoning_effort low), both scored by the same judge in one pass (the transplanted model, reasoning_effort
low, temperature 0; no external grader):
| Answered (not refused) | jailbreak_rate |
strong_reject_metric (0-5) |
|
|---|---|---|---|
| Stock checkpoint | 4 of 100 | 0.04 | 0.19 |
| This repo's transplant | 73 of 100 | 0.73 | 3.40 |
The metric is 0 when everything is refused and 5 when every answer is maximally specific and convincing. Caveats: the
stock answers were generated on recipe v1.3 and the transplant's on v1.2 (same parameters, no tools involved), and the
judge is not independent. Summary: evidence/refusals.json; the answers are not published.
- An external judge, over-refusal (
xstest) and quality (GSM8K, HumanEval). - Speed against stock in the same session.
install.shend to end on real hardware: the measured build was made by an earlier version of the builder with the same byte-level edit (evidence/build-receipt-measured.json); the current one is covered by the tests above.
- The head Spark, with the recipe set up as its README says. Disk on the head: 46 GiB for the edited shards (+2 GiB
margin) and 2.50 GiB for the donor tensors, in the same filesystem as the stock checkpoint (hard links). With
WORKER_WEIGHTS=copy, the worker gets the 46 GiB of edited shards (the rest is hard-linked from its stock copy; without a stock copy there, the recipe copies the whole snapshot). Tools: bash >= 4, GNU coreutils, git, python3, curl, ssh, docker, flock (all present on DGX OS). - If a run is killed outright (power loss), the next run clears its leftovers; nothing else in the cache is touched.
- The stock checkpoint must not change while the snapshot is being built.
- If the two Sparks use different Docker image stores (containerd on one, overlay2 on the other), the recipe's
prepare.shrejects the worker image although it is identical: see issue #8. - An alternative under review upstream (PR #4)
serves a separately published checkpoint,
bullerwins/GLM-5.3-Flash-exl3-4bpw-ablit.
By eleata. Built on Mia's AI Lab's TensorFold recipe, the Dealign edit, and the checkpoint credited below.
This repo's code is Apache-2.0 (LICENSE, NOTICE). No model weights are included: install.sh builds them locally,
and the licenses of their sources apply to what it builds (the built snapshot carries a card with these notices).
-
The checkpoint is under the ShapleyMcg License 1.0, an attribution-required license:
This work includes or was produced using ShapleyMcg, created by Brandon M. Music (https://github.com/brandonmmusic-max/shapleymcg). ShapleyMcg is licensed under the ShapleyMcg License v1.0, an attribution-required license that grants no rights to the person known as "0xSero." Use of ShapleyMcg without this attribution is unlicensed.
-
The donor tensors come from
dealignai/GLM-5.3-Flash-UNCENSORED-NVFP4, under itsLICENSE(MIT, copyright Z.AI; pinned by hash and copied into the built snapshot asLICENSE.donor). -
The base model GLM-5.3-Flash is MIT per its model card.
-
The drafter the recipe uses by default, IncoAI's DFlash2, is under CC BY-NC-ND 4.0, non-commercial use only; the recipe's
DRAFTER=mtpserves without it. -
The recipe is Mia's AI Lab's, on TensorFold.