When quantizing LLMs with a limited precision budget, applying finer quantization globally across all layers outperforms selectively restoring precision to individual critical layers—damage is diffuse, not concentrated.
This paper investigates where quantization damage occurs in large language models and how to best allocate extra precision budget. By systematically testing which layers benefit from higher precision across multiple models, the researchers find that damage is spread across many layers rather than concentrated in specific circuits.