Multiple-choice scoring masks AI model failures in language understanding—QuranicMMLU shows models score 24 percentage points higher on multiple-choice than open-ended tasks, meaning benchmark design significantly impacts what we learn about model capabilities.
QuranicMMLU is a benchmark for testing how well AI models understand Quranic Arabic across five linguistic areas: sound, word structure, grammar, meaning, and context.