12 papers
HandsOnWorld: Unconstrained Egocentric Video Generation with Camera-Disentangled Hand Control
Yushuo Chen, Xiaoyu Shi, Xiaoshi Wu +3
We present HandsOnWorld, a framework for hand-controlled egocentric video generation that learns directly from unconstrained monocular video. Prior generators depend on 3D hand ann…
DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory
Zhenhao Yang, Xiaoshi Wu, Zhengyao Lv +5
Recent advances in video generative models have promoted rapid progress in controllable world models. However, maintaining fine-grained spatio-temporal consistency under long-horiz…
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues
Dingkun Zhou, Krish Patel, Ajay Kankipati +14
Emotions conveyed through voice and face shape engagement and context in human AI interaction. Despite rapid progress in omni modal large language models, the holistic evaluation o…
ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling
Yawen Luo, Xiaoyu Shi, Junhao Zhuang +5
Multi-shot video generation is crucial for long narrative storytelling, yet current bidirectional architectures suffer from limited interactivity and high latency. We propose ShotS…
SemanticGen: Video Generation in Semantic Space
Jianhong Bai, Xiaoshi Wu, Xintao Wang +9
State-of-the-art video generative models typically learn the distribution of video latents in the VAE space and map them to pixels using a VAE decoder. While this approach can gene…
KlingAvatar 2.0 Technical Report
Kling Team, Jialu Chen, Yikang Ding +25
Avatar video generation models have achieved remarkable progress in recent years. However, prior work exhibits limited efficiency in generating long-duration high-resolution videos…