By identifying and prioritizing training on critical reasoning decisions (pivots) rather than all tokens equally, multilingual reasoning transfer becomes more efficient and effective across 17 languages.
This paper improves how large language models learn to reason in multiple languages by focusing training on the most important decision points in reasoning—called 'reasoning pivots'—rather than treating all tokens equally.