Tabular foundation models consistently degrade under distribution shift regardless of their pre-training approach, and practitioners should evaluate OOD robustness before deploying these models in real-world scenarios where data distributions change.
This paper tests nine tabular foundation models on real-world datasets with distribution shifts (like geographic or demographic changes) to see how well they handle data different from their training set. All models degraded under distribution shift, and high-performing models required significant computational resources, raising concerns for deployment in high-stakes applications.