By storing observations in a viewpoint-aware implicit 3D memory rather than explicit depth maps, video world models can generate longer, more consistent videos across different camera angles without running out of token budget.
WorldCrafter is a video world model that maintains consistent 3D-aware memory across different camera viewpoints and long time horizons. It uses an implicit memory system that compresses multi-view observations into tokens optimized for the requested viewpoint, enabling realistic minute-long video generation from a single image or text prompt while respecting what was seen before.