A single transformer block with depth-programmed expert mixtures can replace deep vision encoders, reducing parameters dramatically while maintaining accuracy—useful for efficient vision models and elastic deployment at multiple depths.
This paper proposes reViT, a vision transformer that uses a single block applied repeatedly instead of stacking many layers. By representing the feed-forward network as a mixture of shared experts controlled by depth coordinates, it matches full-depth encoders with 70% fewer parameters and comparable computation. The approach works both for training from scratch and distilling from teacher models.