most citedGenerative Video Compression: Towards 0.01% Compression Rate for Video Transmission

1 citations · 1 across the 25 of their papers we have counts for

collaborators
Showing cs.CVShow all

21 papers · 1 filter

cs.CV2026

Uncertainty DMD: Restoring Diversity in Few-Step Autoregressive Video Distillation

Zixuan Duan, Xunzhi Xiang, Yabo Chen +6

Few-step distillation improves the efficiency of autoregressive (AR) video generation, but often causes diversity collapse: under the same prompt, different noise samples tend to p…

cs.CV2026

Search-to-World: Evaluation of 3D World Delivery from User Request through Web Search

Zixiao Gu, Yabo Chen, Xunzhi Xiang +5

Agentic systems can interpret user requests, search the live web, and use external tools, but their ability to transform retrieved web content into a usable 3D world has not been s…

cs.CV2026

TourPhysics: Bringing Physics to World Models for Exploration and Manipulation from a Single Image

Xin Zhang, Yabo Chen, Zixuan Duan +4

Interactive visual world models must distinguish observation from physical intervention. Camera motion reveals new surfaces, whereas intervention changes object motion, contact, an…

cs.CV2026

RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing

Bojia Zi, Xiaoyan Yang, Yu Zhou +7

Recent advances in video editing have been largely driven by large-scale instruction-based datasets. However, existing datasets still suffer from two critical limitations. First, t…

cs.CV2026

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models

Paribesh Regmi, Qingshuang Chen, Chi Zhang +3

Vision-language models excel at image and video understanding but suffer from high inference latency due to the need to process thousands of tokens per image, limiting their deploy…

cs.CV2026

CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling

Yuyang Huang, Yabo Chen, Wenrui Dai +6

Cinematic video generation is challenging for text-to-video diffusion models due to concurrent requirements on multi-shot generation, fine-grained controllability over characters a…