LLM performance on molecular benchmarks may reflect memorization rather than true predictive capability—models retrieve published values verbatim, and this retrieval is harder to detect and suppress than expected, especially in reasoning-enhanced models.
This paper reveals that frontier language models often retrieve memorized molecular property values verbatim from training data rather than genuinely predicting them. Testing 22 models on 12 benchmarks, researchers found that over 50% of models show exact retrieval on some datasets, and this behavior increases significantly when models use higher reasoning levels.