A compact multimodal model that handles both text and image inputs, producing text output. As a quantized variant of Gemma 4, it trades some precision for reduced memory footprint, making it more accessible on constrained hardware. Its behavior reflects Google's Gemma lineage, repackaged here by coolthor with NVFP4A16 quantization applied.