One representation, three kinds of control. Left: dynamics editing (the dog still pounces on its toy, but the toy is no longer torn apart and stays intact). Middle: training-free retiming (the same run played at 0.5×). Right: re-rendering (the same whisking motion performed by robotic arms). Top row: source video. Bottom row: RVD.
0Method overview
RVD splits a video into two parts: its visual context (the first frame and a caption: what the scene looks like) and a compact dynamics token (how the scene moves and changes over time). This split is learned by self-supervised video reconstruction: a renderer must rebuild the original clip from the first frame plus the token. The first frame already gives the appearance, so the token only has to carry the motion. A language-guided editor can then change this compact token (for example, "cancel the cut" or "slow the run to a walk"), and the renderer turns the edited token back into a video.

1Appearance-controlled re-rendering: comparison with baselines
The dynamics token of the source video is kept fixed; only the visual context (edited first frame and caption) changes. RVD should change the foreground, background, or style while reproducing the source motion and its timing exactly. Watch whether each method keeps the original motion, including hand trajectories, gait phase, and head turns, and whether it stays temporally stable.
2Video dynamics editing: comparison with baselines
Here the instruction asks to change how the scene evolves, not how it looks. RVD edits the dynamics token with the language-guided editor and renders it in the source scene. Appearance editors tend to reproduce the source motion, or they drift in identity and scene.
3More Video Dynamics Edits by RVD
Additional instruction-driven dynamics edits across edit families: state changes (open/closed, cut/whole, broken/intact), interactions (hold/throw), locomotion (walk/run, avoiding an obstacle), and animal behavior (active/resting).
4Training-Free Retiming in Dynamics-Token Space
Because the dynamics token keeps an explicit temporal axis, speed can be changed by resampling the token over time. No editor and no extra training are used, and the renderer is the same frozen one. 0.5× stretches the token (nearest-neighbor resampling followed by light temporal smoothing), so the clip covers only the first half of the original motion. 2× subsamples every other token step, so the full motion is completed in the first half of the clip. All four panels have the same length and play in sync: compare how far the motion has progressed at the same moment.
5More Appearance-Controlled Re-Rendering by RVD
For each source video, the same dynamics token is rendered under three new visual contexts: a new foreground subject, a new background, and a new global style. The motion, including its timing, should match the source in all three.
BibTeX
If you find RVD useful, please cite our work.
@article{Yuan_2026_RVD,
title={{Reimagine Video Dynamics}},
author={Yuan, Yu and Lu, Yawen and Song, Guoxian and Duarte, Kevin and Kalarot, Ratheesh and Chang, Di and Wang, Xijun and Chan, Stanley H.},
journal={arXiv preprint arXiv:},
year={2026}
}