World in World: Explore the World with World Models
Published in arXiv preprint, 2026
World in World is a training-free inference-time interface for controllable exploration with frozen autoregressive video world models. It converts heterogeneous control evidence – source-video observations, target-view scene projections, geometry renderings that complete newly exposed regions, and retrieved generated states beyond the rolling cache – into camera- and time-labelled clean visual states. A correspondence router combines persistent point identities with geometry to establish token correspondences, and evidence-wise attention CFG independently regulates each auxiliary channel using attention responses from the same denoising forward pass. The shared interface supports camera-controlled rerendering, long-horizon revisiting, and human-motion transfer with the same frozen backbone.
Project Page | PDF | Code
Recommended citation: C Song, Y Yang, C Zhang. (2026). "World in World: Explore the World with World Models." arXiv preprint arXiv:2609.11548.
Download Paper