14 papers
LogiShot: Logically Coherent Cross-Shot Video Generation
Shuai Guo, Yuhang Yang, Zeyu Zhang +4
Generating cross-shot videos that are logically connected is essential for content creation. Currently, most cross-shot video-generation workflows, such as short-drama production,…
AffectSeek: Agentic Affective Understanding in Long Videos under Vague User Queries
Zhen Zhang, Yuhang Yang, Yunxiang Jiang +5
Existing affective understanding studies have mainly focused on recognizing emotions from images, audio signals, or pre-cliped video clips, where the affective evidence is already…
Gloria: Consistent Character Video Generation via Content Anchors
Yuhang Yang, Fan Zhang, Huaijin Pi +5
Digital characters are central to modern media, yet generating character videos with long-duration, consistent multi-view appearance and expressive identity remains challenging. Ex…
GIR-Bench: Versatile Benchmark for Generating Images with Reasoning
Hongxiang Li, Yaowei Li, Bin Lin +7
Unified multimodal models integrate the reasoning capacity of large language models with both image understanding and generation, showing great promise for advanced multimodal inte…
End-to-End Spatial-Temporal Transformer for Real-time 4D HOI Reconstruction
Haoyu Zhang, Wei Zhai, Yuhang Yang +2
Monocular 4D human-object interaction (HOI) reconstruction - recovering a moving human and a manipulated object from a single RGB video - remains challenging due to depth ambiguity…
From Frames to Sequences: Temporally Consistent Human-Centric Dense Prediction
Xingyu Miao, Junting Dong, Qin Zhao +3
In this work, we focus on the challenge of temporally consistent human-centric dense prediction across video sequences. Existing models achieve strong per-frame accuracy but often…