4 papers
ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models
Xinye Li, Lingshuai Lin, Lei Wang +8
Action-conditioned video world models require low-latency causal generation and reliable responses to game-native controls. Although causal distillation enables one- or few-step vi…
FlashSign: Pose-Free Guidance for Efficient Sign Language Video Generation
Liuzhou Zhang, Zeyu Zhang, Biao Wu +10
Sign language plays a crucial role in bridging communication gaps between the deaf and hard-of-hearing communities. However, existing sign language video generation models often re…
EgoLCD: Egocentric Video Generation with Long Context Diffusion
Liuzhou Zhang, Jiarui Ye, Yuanlei Wang +6
Generating long, coherent egocentric videos is difficult, as hand-object interactions and procedural tasks require reliable long-term memory. Existing autoregressive models suffer…
VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging
Ming Zhong, Yuanlei Wang, Liuzhou Zhang +6
While Multimodal Large Language Models (MLLMs) excel on benchmarks, their processing paradigm differs from the human ability to integrate visual information. Unlike humans who natu…