A mid-sized multimodal model that handles both text and image inputs, sitting in a practical middle ground within Qwen's third-generation family. It runs in FP8 precision, which reduces memory footprint while preserving much of the full-precision capability — a useful trade-off for those deploying on constrained hardware. Its open-weight Apache 2.0 license makes it freely adaptable for commercial and research use.