Studying how safety alignment is encoded in LLM weights, using Arbitrary-Rank Ablation to isolate and remove the refusal direction in Gemma 4 E2B.
-
Updated
Jul 23, 2026 - Python
Studying how safety alignment is encoded in LLM weights, using Arbitrary-Rank Ablation to isolate and remove the refusal direction in Gemma 4 E2B.
Add a description, image, and links to the arbitrary-rank-ablation topic page so that developers can more easily learn about it.
To associate your repository with the arbitrary-rank-ablation topic, visit your repo's landing page and select "manage topics."