Current multimodal AI models excel at understanding diagrams visually but struggle significantly with converting them to executable code, suggesting this is a critical gap to address for scientific writing tools.
Diagram-MMU is a benchmark with 3.7k scientific diagrams and 18.3k questions that tests how well AI models can understand diagrams and convert them to code. The benchmark evaluates 12 models on three tasks: turning diagrams into LaTeX code, editing diagram code, and answering questions about diagrams—revealing that code generation is much harder for models than visual understanding.