A mid-sized multimodal model that handles both text and image inputs, quantized to int4 precision using AutoRound for reduced memory footprint. The compression makes it more accessible on consumer hardware, though int4 quantization introduces some quality trade-offs compared to full-precision variants. It carries the characteristics of the Qwen3.8 family — reasoning and vision understanding — in a more deployable package.