A mid-sized multimodal model that accepts both text and image inputs, producing text output. It operates as a quantized INT8 W8A16 variant, meaning it trades a small amount of numerical precision for reduced memory footprint and faster inference. Published by lued under an open Apache 2.0 license, it sits in the Qwen3.8 family.