Block-scaled FP4 pretraining can match or exceed current quantization methods while being simpler and faster—no randomized transforms needed, and you can keep FP4 in all layers, not just some.
This paper presents a simpler and more efficient recipe for training large language models using 4-bit floating-point (FP4) precision.