A compact multimodal model that handles both text and images, quantized to 4-bit precision for efficient local deployment. The QAT (Quantization-Aware Training) approach helps preserve capability despite the aggressive compression, making it more reliable than post-hoc quantized alternatives. Runs well on Apple Silicon via the MLX framework, trading some raw fidelity for significantly reduced memory footprint.