A quantized variant of MiniMax's M3 model, optimized for NVFP4 precision to reduce memory footprint while maintaining multimodal capabilities. It accepts both text and image inputs, making it capable of visual understanding tasks alongside text processing. The NVFP4 format suggests it's tuned for efficient inference on NVIDIA hardware, trading some numerical precision for speed and resource efficiency.