A large-scale multimodal model from Qwen's third-generation family, capable of processing both text and images across an exceptionally long context window of one million tokens. Its 125 billion parameters give it substantial capacity for complex reasoning and understanding. It handles long-document and visual tasks within a single context, though as an API-only model, it's not available for local deployment.