A large multimodal model with 26B total parameters but only 4B active at inference time, quantized to FP8 dynamic for efficient deployment. It accepts both text and image inputs and produces text outputs. As a diffusion-based language model variant, its generation behavior may differ from standard autoregressive models, though specific capability details are limited.