A quantized variant of Qwen3.5 9B that trades a small amount of numerical precision for reduced memory footprint and faster inference. The MXFP8 weight quantization combined with FP8 key-value cache and FP8 attention means nearly every compute-heavy operation runs in lower precision, calibrated via AutoRound. Expect behavior very close to the base model, though edge cases involving subtle numerical distinctions may occasionally differ.