A technique for extending pretrained models to multi-view or multi-modal outputs by tiling latent representations spatially without modifying the core architecture.