Official implementation of The Generalization Ridge: Information Flow in Natural Language Generation, accepted to COLM 2026.
We propose InfoRidge, an information-theoretic framework, to characterize how predictive information — the mutual information between hidden representations and target outputs — varies across depth during training. Our experiments across various models and datasets reveal a consistent non-monotonic trend: predictive information peaks in intermediate layers — forming a generalization ridge — before declining in final layers, reflecting a transition between generalization and memorization.
conda create -n inforidge python=3.12.3
conda activate inforidge
pip install -r requirements.txtfinetune/
train.py
mutual_information/
main.py
data_loader.py
embedding.py
analysis.py
utils.py
multiple/
residual/
fit_beta.py
scripts/
1_finetune.sh
2_mutual_information.sh
3_multitoken.sh
4_residual_beta.sh
CKPT_ROOT=./checkpoints bash scripts/1_finetune.sh
CKPT_ROOT=./checkpoints bash scripts/2_mutual_information.sh
bash scripts/3_multitoken.sh
CKPT_ROOT=./checkpoints bash scripts/4_residual_beta.shEach defaults to the ECQA / Qwen2.5-0.5B setting. See the paper for the full set of configurations it reports.
Fine-tune a pretrained language model on one of the downstream tasks.
python finetune/train.py \
--dataset ecqa --model_name Qwen/Qwen2.5-0.5B \
--output_dir checkpoints/ecqa_qwen --epochs 3 --save_steps 100 --seed 42--dataset is synthetic, clutrr or ecqa.
We use the public CLUTRR (Sinha et al., 2019) and ECQA (Aggarwal et al., 2021) datasets. The synthetic arithmetic dataset is generated directly by our code.
Compute predictive information between the hidden representation at each layer and the target.
python mutual_information/main.py \
--checkpoint_base checkpoints/ecqa_qwen --dataset_type ecqa \
--sample_count 10000 --seeds 42 --output_dir results/izy/ecqa_qwenFor each layer, we construct Gaussian-kernel Gram matrices from the sampled hidden representations Z and target representations Y. We then compute their matrix-based entropies from the eigenvalue spectra of the normalized Gram matrices and estimate predictive information as
I = H(Z) + H(Y) − H(Z,Y),
where the joint Gram matrix is obtained from the Hadamard product of the Gram matrices for Z and Y.
Evaluate the layer-wise information profile throughout autoregressive generation.
python mutual_information/multiple/main.py \
--model_path meta-llama/Llama-3.1-8B --dataset_type cnn_dailymail \
--sample_count 100 --max_new_tokens 50 --mini_batch_size 3 \
--temperature 0.7 --out_dir results/multitoken/llamaWe use CNN/DailyMail (Hermann et al., 2015; See et al., 2017) for this experiment.
Learn one residual scaling coefficient per transformer block to measure layer-wise reliance under different data regimes.
python residual/fit_beta.py \
--checkpoint_dir checkpoints/ecqa_qwen/checkpoint-1425 \
--model_type qwen --output_dir results/residual/ecqa_qwenFreezes the model and learns one scalar β per block, rescaling its residual
contribution as y' = x + β(y − x).
@inproceedings{chang2026generalization,
title = {The Generalization Ridge: Information Flow in Natural Language Generation},
author = {Chang, Ruidi and Deng, Chunyuan and Chen, Hanjie},
booktitle = {Conference on Language Modeling (COLM)},
year = {2026},
eprint = {2507.05387},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2507.05387}
}