Mage VL is a multimodal model from Microsoft that accepts both text and image inputs to produce text outputs. It operates under an open Apache 2.0 license, making its weights freely available in safetensors format. Details about its specific capabilities and performance characteristics are limited beyond its vision-language architecture.