You can train one model that works well at any depth by supervising random layer prefixes during training—no architectural tricks needed, just a smarter training objective that makes the model elastic across compute budgets.
This paper introduces Telescopic Language Models (TLMs), which train a single model that works effectively at every layer depth rather than requiring separate models for different compute budgets.