Vision-language models can extract simple database schema elements from diagrams but fail on advanced ER constructs—this benchmark reveals critical gaps in multimodal understanding of structured technical diagrams that matter for AI-assisted database engineering.
ERUnderstand is a benchmark dataset of 2,960 Entity-Relationship diagrams that tests how well AI vision-language models can understand database schemas from images. The researchers found that while models handle basic diagram elements well, they struggle significantly with complex features like weak entities and N-ary relationships, even when augmented with reasoning capabilities.