A mid-sized multimodal model that handles both text and image inputs, quantized to 8-bit for more efficient memory usage. The MLX-community packaging makes it particularly suited for running locally on Apple Silicon hardware. The 8-bit quantization means a trade-off: reduced memory footprint compared to full precision, with some potential impact on output fidelity.