A compact, quantized vision-language model that trades some precision for dramatically reduced memory footprint through NVFP4 quantization. It handles both text and image inputs, making it capable of multimodal tasks in resource-constrained environments. The quantization approach means it runs faster and leaner than full-precision counterparts, though with potential minor quality trade-offs.