Randomly removing input tokens during training to reduce computational cost while improving model robustness.