[ICLR 2025] When Attention Sink Emerges in Language Models: An Empirical View (Spotlight)
-
Updated
Jul 8, 2025 - Python
[ICLR 2025] When Attention Sink Emerges in Language Models: An Empirical View (Spotlight)
Official repository for RegToken - ECCV 2026
🐙 Implements Flash Attention with sink for gpt-oss-20b; includes test.py. WIP backward pass, varlen support, and community sync to return softmax_lse only.
Effective rank, RankMe, E1, CKA and anisotropy on transformer hidden states are determined by one direction. The exact identity, and the attention sink behind it.
Benchmark of KV-cache eviction policies on GPT-2-medium showing two clear regimes: sink-preserving heuristics dominate at tiny budgets, while attention-based eviction becomes near-lossless at moderate budgets. attention_384 matches full-cache quality with ~33% less cache.
Empirical comparison of Absolute PE vs RoPE showing that absolute PE is 63× more sensitive to sink token removal than RoPE. Includes attention sink measurement, KV-cache eviction tolerance, masked-key ablation, perplexity scaling, and decode throughput.
Empirical measurement of attention sink in GPT-2: attention map analysis, per-head classification, and masked-key ablation to assess functional impact on tail perplexity and output distribution.
Add a description, image, and links to the attention-sink topic page so that developers can more easily learn about it.
To associate your repository with the attention-sink topic, visit your repo's landing page and select "manage topics."