1 citations · 1 across the 25 of their papers we have counts for
21 papers · 1 filter
Uncertainty DMD: Restoring Diversity in Few-Step Autoregressive Video Distillation
Zixuan Duan, Xunzhi Xiang, Yabo Chen +6
Few-step distillation improves the efficiency of autoregressive (AR) video generation, but often causes diversity collapse: under the same prompt, different noise samples tend to p…
Search-to-World: Evaluation of 3D World Delivery from User Request through Web Search
Zixiao Gu, Yabo Chen, Xunzhi Xiang +5
Agentic systems can interpret user requests, search the live web, and use external tools, but their ability to transform retrieved web content into a usable 3D world has not been s…
TourPhysics: Bringing Physics to World Models for Exploration and Manipulation from a Single Image
Xin Zhang, Yabo Chen, Zixuan Duan +4
Interactive visual world models must distinguish observation from physical intervention. Camera motion reveals new surfaces, whereas intervention changes object motion, contact, an…
RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing
Bojia Zi, Xiaoyan Yang, Yu Zhou +7
Recent advances in video editing have been largely driven by large-scale instruction-based datasets. However, existing datasets still suffer from two critical limitations. First, t…
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models
Paribesh Regmi, Qingshuang Chen, Chi Zhang +3
Vision-language models excel at image and video understanding but suffer from high inference latency due to the need to process thousands of tokens per image, limiting their deploy…
CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling
Yuyang Huang, Yabo Chen, Wenrui Dai +6
Cinematic video generation is challenging for text-to-video diffusion models due to concurrent requirements on multi-shot generation, fine-grained controllability over characters a…