A diffusion model trained on multiple views of scenes to understand 3D structure and generate consistent views.