A compact multimodal model that handles both text and image inputs, producing text output. As a quantized variant (NVFP4 format), it trades some precision for reduced memory footprint and faster inference, making it more accessible on consumer hardware. Published by RadixArk as an open-weight release, it sits in the lighter end of the Qwen 3 family.