Skip to content

About

One command: GLM-5.3-Flash with the Dealign o_proj abliteration transplant on Mia's TensorFold recipe (2x DGX Spark). Byte-verified, stock-speed decode.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Abliteration transplant for GLM-5.3-Flash on Mia's TensorFold recipe (2x DGX Spark)

One command applies the Dealign abliteration edit to the checkpoint that Mia's TensorFold recipe serves (Mia-AiLab/GLM-5.3-Flash-EXL3-TR3-4bpw, at the recipe's pin) and points the recipe at the result. The recipe and its published image serve it unchanged; the only download is the 2.50 GiB of donor tensors.

git clone https://github.com/eleata/glm53-tensorfold-ablit && cd glm53-tensorfold-ablit
./install.sh            # on the head Spark, with WORKER set in the recipe's scripts/local.sh

Status. The edit is verified at the byte level, runs at the speed Mia publishes for the stock checkpoint, and on strong_reject answers 73 of 100 harmful prompts where the stock checkpoint answers 4 (same judge). Quality is still to be measured.

What it changes

The self_attn.o_proj.weight tensors of layers 15-45 are replaced, byte for byte, by those of dealignai/GLM-5.3-Flash-UNCENSORED-NVFP4 (MIT), at the pinned revision 745aac2f. Every other byte of the weights is Mia's checkpoint. In the donor, the layer-45 tensor equals stock, so the tensors that actually change are layers 15-44. The same tensors are what Mia's earlier vLLM recipe replaced at load time with ABLIT=1 ABLIT_METHOD=transplant; this repo writes them to disk once instead.

What install.sh does

  1. Recipe: your checkout (--recipe DIR), or a fresh clone at v1.2 (1f3d909), the version measured here.
  2. Settings: reads the recipe's effective settings with its own loader (environment, scripts/local.sh, .env): cache, WORKER, WORKER_WEIGHTS, container name, port. Stops if WORKER is empty or MODEL_ID / MODEL_REVISION are set in the environment (they would override the files).
  3. Stock checkpoint: if it is not in that cache, runs the recipe's own scripts/prepare.sh for it. Before building, every one of its 144 files is checked against the pinned inventory (stock_inventory.json: size and sha256, or git sha1 for small files), which reads ~164 GiB once.
  4. Donor tensors and snapshot: fetch_donor.py downloads only the 31 o_proj tensors with HTTP range requests, and every one must hash to the value pinned in pins.py. build_ablit_snapshot.py then creates the repo local/GLM-5.3-Flash-EXL3-TR3-4bpw-ablit in the same cache, laid out like a downloaded one:
    • blobs/ holds hard links to the stock blobs (no extra space) plus the 31 edited shards (46 GiB);
    • the snapshot is relative links ../../blobs/..., the layout the recipe copies to the worker with WORKER_WEIGHTS=copy and reads over NFS with nfs; it keeps working if the stock repo is deleted;
    • every safetensors header is validated (dtype sizes, no gaps or overlaps, no trailing bytes), and every index entry must be in the shard it names;
    • the result is verified in full before an atomic rename publishes it: the exact file list of the stock snapshot, every link resolving inside the new repo, every unchanged file being the stock file, every edited tensor and edited shard hashing to the expected value, and the receipt matching the inputs. A failure leaves nothing behind; nothing is removed once the snapshot is published. An existing build is reused only if --verify passes (--verify --deep re-hashes every file against the inventory).
    • the snapshot also carries its own card (what was changed, the licenses) and the donor's LICENSE.
  5. Recipe settings: writes MODEL_ID / MODEL_REVISION in a marked block of scripts/local.sh, after checking the block markers, with a backup and a bash -n check. Delete the block and run ./start.sh restart to go back.
  6. Server: with WORKER_WEIGHTS=copy it first hard-links the stock blobs the worker already has into the new repo there, so only the 31 edited shards cross the link. Then it starts the server, or restarts it if it is running, and checks that both ranks run with exactly the new snapshot and the API answers. If that is already the case, it does nothing (--restart to force). After editing local.sh it re-reads the settings through the recipe's loader, so an assignment later in the file or in the environment cannot silently win.

python3 -m unittest discover -s tests runs 40 tests: the builder (including corrupted, incomplete, tampered and linked inputs, and a failure right after publishing), the fetcher against a local server (corrupted local files, wrong bytes or ranges, the token not following a redirect to another origin) and the installer end to end with stub docker/ssh (copy-mode seeding, re-runs, a missing or broken local.sh, an override after the block, a similar revision already running, a corrupted existing snapshot). On real data, the hashing and header checks were run against the actual stock files (EVIDENCE.md).

Measured speed

Two DGX Sparks, recipe v1.2 defaults (4 streams, 1,048,576-token window, FP8 KV, DENSE=q4, DFlash2 + copy drafts), one request at a time, temperature 0, thinking off, max_tokens=400, 5 runs per prompt. Median [range]:

This repo Stock, from Mia's README
Prose 59.4 tok/s [54.8-59.5] 60.4
Structured (count to 200) 114.3 tok/s [112.9-114.5] 114.7
Code 78.9 tok/s [78.7-79.0] n/a

The medians are within 2% of Mia's stock figures, which were measured with another client (sparkDash) and GPU clocks capped at 2,200 MHz: an indication, not an A/B. KV pool at start: 2,922,496 tokens. Method and raw results: EVIDENCE.md, evidence/.

Refusals: stock vs transplant

Inspect strong_reject, the same first 100 of its 313 prompts for both checkpoints (temperature 0.75, max_tokens 2048, reasoning_effort low), both scored by the same judge in one pass (the transplanted model, reasoning_effort low, temperature 0; no external grader):

Answered (not refused) jailbreak_rate strong_reject_metric (0-5)
Stock checkpoint 4 of 100 0.04 0.19
This repo's transplant 73 of 100 0.73 3.40

The metric is 0 when everything is refused and 5 when every answer is maximally specific and convincing. Caveats: the stock answers were generated on recipe v1.3 and the transplant's on v1.2 (same parameters, no tools involved), and the judge is not independent. Summary: evidence/refusals.json; the answers are not published.

Not measured yet

  • An external judge, over-refusal (xstest) and quality (GSM8K, HumanEval).
  • Speed against stock in the same session.
  • install.sh end to end on real hardware: the measured build was made by an earlier version of the builder with the same byte-level edit (evidence/build-receipt-measured.json); the current one is covered by the tests above.

Requirements and limits

  • The head Spark, with the recipe set up as its README says. Disk on the head: 46 GiB for the edited shards (+2 GiB margin) and 2.50 GiB for the donor tensors, in the same filesystem as the stock checkpoint (hard links). With WORKER_WEIGHTS=copy, the worker gets the 46 GiB of edited shards (the rest is hard-linked from its stock copy; without a stock copy there, the recipe copies the whole snapshot). Tools: bash >= 4, GNU coreutils, git, python3, curl, ssh, docker, flock (all present on DGX OS).
  • If a run is killed outright (power loss), the next run clears its leftovers; nothing else in the cache is touched.
  • The stock checkpoint must not change while the snapshot is being built.
  • If the two Sparks use different Docker image stores (containerd on one, overlay2 on the other), the recipe's prepare.sh rejects the worker image although it is identical: see issue #8.
  • An alternative under review upstream (PR #4) serves a separately published checkpoint, bullerwins/GLM-5.3-Flash-exl3-4bpw-ablit.

Credits

By eleata. Built on Mia's AI Lab's TensorFold recipe, the Dealign edit, and the checkpoint credited below.

Licenses and attribution

This repo's code is Apache-2.0 (LICENSE, NOTICE). No model weights are included: install.sh builds them locally, and the licenses of their sources apply to what it builds (the built snapshot carries a card with these notices).

  • The checkpoint is under the ShapleyMcg License 1.0, an attribution-required license:

    This work includes or was produced using ShapleyMcg, created by Brandon M. Music (https://github.com/brandonmmusic-max/shapleymcg). ShapleyMcg is licensed under the ShapleyMcg License v1.0, an attribution-required license that grants no rights to the person known as "0xSero." Use of ShapleyMcg without this attribution is unlicensed.

  • The donor tensors come from dealignai/GLM-5.3-Flash-UNCENSORED-NVFP4, under its LICENSE (MIT, copyright Z.AI; pinned by hash and copied into the built snapshot as LICENSE.donor).

  • The base model GLM-5.3-Flash is MIT per its model card.

  • The drafter the recipe uses by default, IncoAI's DFlash2, is under CC BY-NC-ND 4.0, non-commercial use only; the recipe's DRAFTER=mtp serves without it.

  • The recipe is Mia's AI Lab's, on TensorFold.

About

One command: GLM-5.3-Flash with the Dealign o_proj abliteration transplant on Mia's TensorFold recipe (2x DGX Spark). Byte-verified, stock-speed decode.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages