By adding recurrent feedback during pretraining via teacher-supervised state prediction, language models can improve reasoning and task performance without sacrificing training efficiency, suggesting that feed-forward architectures unnecessarily limit information flow.
This paper introduces LIFT, a transformer architecture that enables information to flow backward across layers during language model generation. Instead of the standard feed-forward design, LIFT uses teacher supervision during pretraining to train models to predict both the next token and a dense state representation derived from a teacher model.