Post-training with domain-specific data, reasoning traces, and test-time compute strategies can push language models to superhuman performance on complex reasoning tasks like competitive programming.
Researchers trained specialized language models to excel at competitive programming by combining curated problem datasets, synthetic reasoning traces, fine-tuning, and reinforcement learning. Their system achieved gold-medal performance on the International Olympiad in Informatics (IOI) 2025 and 2026, becoming the first AI to outscore the top human competitor on an IOI problem set.