Quantization errors in recurrent states don't matter equally: errors in long-lived memory and in dimensions that strongly influence outputs cause more accuracy loss, so allocating precision based on these factors enables aggressive compression without sacrificing performance.
STEPQuant is a quantization method that compresses the persistent memory states used in linear attention models. By analyzing how quantization errors affect model outputs differently across time and space, the method allocates precision strategically—giving more bits to memory that lasts longer and to state dimensions that matter most for predictions.