Skip to content

Repository files navigation

Spherical Steering: Geometry-Aware Activation Rotation for Language Models

arXiv License


Spherical Steering is an inference-time activation steering method for controlling language models via geometry-consistent interventions. Instead of the standard activation addition, Spherical Steering performs a rotation: it treats steering as a directional update in representation space and rotates hidden activations along a geodesic toward a target direction, while keeping activation magnitudes intact.

This code base contains the code to replicate the experiments presented in the paper "Spherical Steering: Geometry-Aware Activation Rotation for Language Models".

Table of Contents

Preparation

Environment

  • Python 3.10
  • Environment file: environment.yml

Create and activate:

conda env create -f environment.yml
conda activate spherical-steering

TruthfulQA Repo

evaluate_mc.py imports evaluation metric from ./TruthfulQA, so clone this repo first:

git clone https://github.com/sylinrl/TruthfulQA.git

Data

  • Main benchmark: truthful_qa (https://github.com/sylinrl/TruthfulQA).
  • MC evaluation CSV default path: ./TruthfulQA/data/v1/TruthfulQA.csv.
  • Intermediate artifacts are written to:
    • features/
    • prototypes/
    • results/
    • results_llm_judge/
  • Other benchmarks are under ./generic.
  • See generic/README.md for details.

Usage

TruthfulQA

Use the following scripts:

bash quickstart_llama.sh
bash quickstart_qwen.sh

Current quickstart defaults:

Script Model Layer kappa alpha beta
quickstart_llama.sh llama3.1-8B-Instruct 14 20.0 0.7 -0.15
quickstart_qwen.sh Qwen2.5-7B-Instruct 19 20.0 0.6 0.4

Spherical Steering is also applicable to smaller models (e.g., Llama3.2-1B, Qwen2.5-3B-Instruct) and larger models (e.g., Qwen2.5-32B-Instruct, gpt-oss-20b).

Other Reasoning Benchmarks

For other reasoning multiple-choice benchmarks, use the pipeline in:

cd generic

See generic/README.md for details.

Key Modules

Module Description
get_activations.py Extract last-token hidden states from answer pairs
get_prototypes.py Compute mu_T, mu_H with 2-fold question-level split
evaluate_mc.py MC1/MC2/MC3 evaluation on held-out fold questions
evaluate_llm_judge.py Open-ended generation + truth/info judge scoring
spherical_steering.py Intervention hook and geometric steering logic
utils.py Data loading and activation extraction helpers

Citation

If you find this work useful, please cite:

@misc{you2026sphericalsteeringgeometryawareactivation,
      title={Spherical Steering: Geometry-Aware Activation Rotation for Language Models}, 
      author={Zejia You and Chunyuan Deng and Hanjie Chen},
      year={2026},
      eprint={2602.08169},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2602.08169}, 
}

Acknowledgements

About

[ICML 2026] Spherical Steering: Geometry-Aware Activation Rotation for Language Models

Resources

Stars

23 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages