A compact multimodal model that handles both text and image inputs, producing text outputs. It carries an instruction-tuned disposition, meaning it responds to direct prompts and task-oriented requests rather than operating as a raw completion engine. As an open-weight release under Apache 2.0, it can be run and modified locally, which suits teams who need control over deployment.