An NVFP4-quantized variant of Zhipu AI's GLM-5.2 model, compressed to 4-bit precision for efficient deployment on NVIDIA hardware. Retains the base model's 1M-token context window and general-purpose text capabilities while significantly reducing memory requirements. The quantization trade-off is minimal for most tasks, making this a practical choice for running a large-context model on constrained GPU setups.