12 papers
DRM: Diffusion-based Reward Model With Step-wise Guidance
Jaxon Zhang, Binxin Yang, Hubery Yin +2
Current mainstream methods of aligning diffusion models with human preferences typically employ VLM-based reward models. However, these reward models, pre-trained for semantic alig…
OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation
Lei Zhu, Xing Cai, Yingjie Chen +6
Recent advancements in audio-video joint generation models have demonstrated impressive capabilities in content creation. However, generating high-fidelity human-centric videos in…
VersusQ: Pairwise Margin Reasoning for Generalizable Video Quality Assessment
Shibei Meng, Binxin Yang, Yuan Liu +4
Large Multimodal Models (LMMs) have shown promise for video quality assessment, but most methods still predict an absolute score for each video. Such pointwise supervision often mi…
Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing
Zeyue Tian, Binxin Yang, Zhaoyang Liu +8
Recent progress in multimodal models has spurred rapid advances in audio understanding, generation, and editing. However, these capabilities are typically addressed by specialized…
Identity as Presence: Towards Appearance and Voice Personalized Joint Audio-Video Generation
Qin Chen, Yingjie Chen, Shilun Lin +9
Recent advances in video synthesis have enabled realistic integration of real individuals, driving demand for identity-aware generation. While emerging methods support joint appear…
NOVA: Sparse Control, Dense Synthesis for Pair-Free Video Editing
Tianlin Pan, Jiayi Dai, Chenpu Yuan +7
Recent video editing models have achieved impressive results, but most still require large-scale paired datasets. Collecting such naturally aligned pairs at scale remains highly ch…