A compact multimodal model that handles both text and image inputs, sitting in the lighter end of the Qwen 3 family. The FP8 quantization means it trades a small amount of numerical precision for reduced memory footprint and faster inference, making it more practical to run on constrained hardware. Details about its specific capability profile beyond multimodal input handling are limited.