A large multimodal model that accepts both text and image inputs, producing text outputs. It runs in FP8 dynamic quantization, meaning it trades a small amount of numerical precision for significantly reduced memory footprint and faster inference — a practical choice when deploying at scale. The open-weight release via EmbeddedLLM makes it accessible for self-hosted setups.