ALiBi positional encoding silently fails on long sequences due to numerical precision issues, degrading retrieval performance; using log-scaled distances instead of linear scaling provides the most reliable fix.
This paper discovers that ALiBi positional encoding—a popular method for handling long sequences in transformers—has a critical flaw: its linear bias scaling causes floating-point underflow, making many attention weights zero and blinding attention heads. The authors analyze this problem, test fixes, and show it hurts token retrieval tasks while barely affecting standard benchmarks.