Spherical Steering is an inference-time activation steering method for controlling language models via geometry-consistent interventions. Instead of the standard activation addition, Spherical Steering performs a rotation: it treats steering as a directional update in representation space and rotates hidden activations along a geodesic toward a target direction, while keeping activation magnitudes intact.
This code base contains the code to replicate the experiments presented in the paper "Spherical Steering: Geometry-Aware Activation Rotation for Language Models".
- Python
3.10 - Environment file:
environment.yml
Create and activate:
conda env create -f environment.yml
conda activate spherical-steeringevaluate_mc.py imports evaluation metric from ./TruthfulQA, so clone this repo first:
git clone https://github.com/sylinrl/TruthfulQA.git- Main benchmark:
truthful_qa(https://github.com/sylinrl/TruthfulQA). - MC evaluation CSV default path:
./TruthfulQA/data/v1/TruthfulQA.csv. - Intermediate artifacts are written to:
features/prototypes/results/results_llm_judge/
- Other benchmarks are under
./generic. - See
generic/README.mdfor details.
Use the following scripts:
bash quickstart_llama.sh
bash quickstart_qwen.shCurrent quickstart defaults:
| Script | Model | Layer | kappa |
alpha |
beta |
|---|---|---|---|---|---|
quickstart_llama.sh |
llama3.1-8B-Instruct |
14 | 20.0 | 0.7 | -0.15 |
quickstart_qwen.sh |
Qwen2.5-7B-Instruct |
19 | 20.0 | 0.6 | 0.4 |
Spherical Steering is also applicable to smaller models (e.g., Llama3.2-1B, Qwen2.5-3B-Instruct) and larger models (e.g., Qwen2.5-32B-Instruct, gpt-oss-20b).
For other reasoning multiple-choice benchmarks, use the pipeline in:
cd genericSee generic/README.md for details.
| Module | Description |
|---|---|
get_activations.py |
Extract last-token hidden states from answer pairs |
get_prototypes.py |
Compute mu_T, mu_H with 2-fold question-level split |
evaluate_mc.py |
MC1/MC2/MC3 evaluation on held-out fold questions |
evaluate_llm_judge.py |
Open-ended generation + truth/info judge scoring |
spherical_steering.py |
Intervention hook and geometric steering logic |
utils.py |
Data loading and activation extraction helpers |
If you find this work useful, please cite:
@misc{you2026sphericalsteeringgeometryawareactivation,
title={Spherical Steering: Geometry-Aware Activation Rotation for Language Models},
author={Zejia You and Chunyuan Deng and Hanjie Chen},
year={2026},
eprint={2602.08169},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2602.08169},
}