Looping computation within denoising steps is more efficient than scaling model size—you can get better images with fewer parameters by iteratively refining representations through repeated block execution.
This paper proposes Looped Diffusion Transformer, which improves text-to-image generation by repeatedly processing the same Transformer blocks within each denoising step rather than making models larger.