5 papers · 1 filter
EraseSAE: Surgical Concept Erasure in Text-to-Video Diffusion Models via Sparse Autoencoders
Xinghao Wang, Dong Li, Wei Yu +5
Recent advances in text-to-video (T2V) diffusion models have demonstrated remarkable generative capabilities, yet their reliance on loosely curated training data raises pressing sa…
Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory
Runjia Qian, Zile Wang, Jihai Zhang +14
Interactive world models extend video generation from offline clip synthesis toward persistent simulation of interactive virtual worlds, enabling applications in games, robotics, e…
EmambaIR: Efficient Visual State Space Model for Event-guided Image Reconstruction
Wei Yu, Yunhang Qian
Recent event-based image reconstruction methods predominantly rely on Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) to process complementary event information…
MosaicMem: Hybrid Spatial Memory for Controllable Video World Models
Wei Yu, Runjia Qian, Yumeng Li +8
Video diffusion models are moving beyond short, plausible clips toward world simulators that must remain consistent under camera motion, revisits, and intervention. Yet spatial mem…
Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose Interaction
Dong Li, Wenqi Zhong, Wei Yu +5
Video virtual try-on aims to seamlessly dress a subject in a video with a specific garment. The primary challenge involves preserving the visual authenticity of the garment while d…