🫥 ReMind — video generation with dynamic memory

ReMind is a Wan2.2-TI2V-5B world model taught to remember what it can no longer see. Schedule a disturbance mid-clip — an object covers the lens, the lights go out, the camera turns away — and the scene keeps evolving out of sight, so when the view comes back the state has moved on instead of resetting.

81 frames · 832×480 · 16 fps · four-step DMD rollout over seven 3-frame chunks.

Out-of-sight event
Scheduled over chunks 3–5 of 7, with two recovery chunks after it.

Examples from the ReMind project page

Image-to-video (two occluder-recovery cases, three clean)
Camera-controlled (paired trajectories, PM-RoPE)

The occluder and the lights-out event are never painted into the pixels — they are described to the generator through chunk-local text, exactly as in the paper. The camera mode instead feeds a real trajectory (extrinsics + intrinsics) into ReMind's camera-phase RoPE; leave the trajectory file empty to use a synthetic pan-away-and-return.

📄 Paper · 💻 Code · 🤗 Weights (CC-BY-NC-4.0, research use) · base model Wan2.2-TI2V-5B