A multimodal model that accepts both text and image inputs, producing text outputs. It operates at 4-bit quantization, which reduces memory footprint at the cost of some precision compared to full-precision variants. As an open-weight release from mlx-community, it fits into workflows where local deployment and accessibility matter.