🫥 ReMind — video generation with dynamic memory
ReMind is a Wan2.2-TI2V-5B world model taught to remember what it can no longer see. Schedule a disturbance mid-clip — an object covers the lens, the lights go out, the camera turns away — and the scene keeps evolving out of sight, so when the view comes back the state has moved on instead of resetting.
81 frames · 832×480 · 16 fps · four-step DMD rollout over seven 3-frame chunks.
Examples from the ReMind project page
The occluder and the lights-out event are never painted into the pixels — they are described to the generator through chunk-local text, exactly as in the paper. The camera mode instead feeds a real trajectory (extrinsics + intrinsics) into ReMind's camera-phase RoPE; leave the trajectory file empty to use a synthetic pan-away-and-return.
📄 Paper · 💻 Code · 🤗 Weights (CC-BY-NC-4.0, research use) · base model Wan2.2-TI2V-5B