A compact multimodal model that handles both text and image inputs, producing text responses. At 8B parameters, it sits in the efficient tier of vision-language models, balancing capability with accessibility. It follows an instruction-tuned format, meaning it's designed to respond to direct prompts rather than operate as a raw base model.