A compact multimodal model from the Qwen 3 family, fine-tuned and republished by jcbtc. It accepts both text and image inputs, making it capable of visual understanding tasks alongside text processing. As a smaller model in the family, it trades raw capacity for accessibility and efficiency.