Yu Yuan · Research overview
Towards physics-consistent andcontrollable generative worlds.
Generative models should model the physics of worlds—not just pixels.
Navigate with ← →, Space, or swipe · M for map · F for full screen
Can generative models follow physical rules across camera, dynamics, space–time, and action?
01 / 04
CVPR 2025 HighlightGenerative Photography: Scene-Consistent Camera Control for Realistic Text-to-Image Synthesis
Control the camera, preserve the world.
Scene-consistent camera control for realistic text-to-image synthesis.
ProblemExisting generative models do not understand the physical meaning of camera intrinsics or preserve scene consistency.
IdeaDimensional lifting disentangles scene and camera; differentiated parameter injection enables fine-grained control.
ResultControl bokeh, focal length, shutter speed, and color temperature while preserving scene identity.
02 / 04
ICLR 2026NewtonGen: Physics-Consistent and Controllable Text-to-Video Generation via Neural Newtonian Dynamics
Give video generation a sense of physics.
Physics-consistent and controllable text-to-video generation.
ProblemVideo models lack physical consistency and parameter-level controllability.
IdeaLearn Neural Newtonian Dynamics, then guide video generation with explicit physical states.
ResultDiverse, physically consistent, and controllable video generation.
03 / 04
CVPR 2026SeeU: Seeing the Unseen World via 4D Dynamics-Aware Generation
See beyond the observed time and space.
Continuous 4D dynamics modeling and video generation.
ProblemInfer the unobserved world from sparse 2D observations.
IdeaLearn continuous 4D dynamics as the backbone of 4D world generation.
ResultTemporal, spatial, and edited unseen-world generation.
04 / 04
Under ReviewOptiWorld: Optimal Control for Video World Generation under Physical Constraints
Plan a better future, then render it.
OptiWorld makes video world models choose better futures: safer, smoother, more efficient, and physically constrained motions before rendering.
ProblemCurrent video generation lacks proactive understanding and control, producing results that may be unsafe or inefficient.
IdeaIntroduce classical optimal control before video rendering.
ResultSafer, more efficient goal-conditioned I2V and video refinement.