A mid-sized multimodal model that handles both text and image inputs, quantized to NVFP4 precision for efficient inference on RTX 5090 hardware. The NVFP4 format trades some numerical precision for significantly reduced memory footprint and faster throughput on compatible GPUs. As an open-weight release under Apache 2.0, it sits in the practical deployment tier — capable enough for real workloads, optimized for specific consumer hardware.