A quantization format where weights are compressed to 4-bit precision while activations remain at 16-bit precision, reducing memory usage while maintaining reasonable accuracy.