OvisOCR2 is a multimodal model built on the Qwen3.5 family, designed to process both images and text as inputs and return text outputs. Its image-plus-text input combination positions it for document understanding and visual text extraction tasks. As an open-weight model under Apache 2.0, it can be run and modified locally without licensing restrictions.