A multimodal open-weight model from DeepSeek that accepts both text and image inputs. It carries the MIT license, making weights freely usable and modifiable. Beyond its multimodal input support, specific capability details for this particular version are limited in available documentation.