A compact multimodal model that handles both text and image inputs, quantized to NVFP4 format by Unsloth for efficient deployment. The reduced precision keeps memory footprint small while aiming to preserve the underlying Gemma 4 capabilities. Trade-offs typical of aggressive quantization may appear in nuanced reasoning tasks.