activity
20242026
most citedAudio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing

1 citations · 1 across the 5 of their papers we have counts for

collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV2026

DRM: Diffusion-based Reward Model With Step-wise Guidance

Jaxon Zhang, Binxin Yang, Hubery Yin +2

Current mainstream methods of aligning diffusion models with human preferences typically employ VLM-based reward models. However, these reward models, pre-trained for semantic alig…

cs.CV2026

OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation

Lei Zhu, Xing Cai, Yingjie Chen +6

Recent advancements in audio-video joint generation models have demonstrated impressive capabilities in content creation. However, generating high-fidelity human-centric videos in…

cs.CV2026

VersusQ: Pairwise Margin Reasoning for Generalizable Video Quality Assessment

Shibei Meng, Binxin Yang, Yuan Liu +4

Large Multimodal Models (LMMs) have shown promise for video quality assessment, but most methods still predict an absolute score for each video. Such pointwise supervision often mi…

cs.CV2026

Identity as Presence: Towards Appearance and Voice Personalized Joint Audio-Video Generation

Qin Chen, Yingjie Chen, Shilun Lin +9

Recent advances in video synthesis have enabled realistic integration of real individuals, driving demand for identity-aware generation. While emerging methods support joint appear…

cs.CV2026

NOVA: Sparse Control, Dense Synthesis for Pair-Free Video Editing

Tianlin Pan, Jiayi Dai, Chenpu Yuan +7

Recent video editing models have achieved impressive results, but most still require large-scale paired datasets. Collecting such naturally aligned pairs at scale remains highly ch…

cs.CV2026

StableWorld: Towards Stable and Consistent Long Interactive Video Generation

Ying Yang, Zhengyao Lv, Tianlin Pan +6

In this paper, we explore the overlooked challenge of stability and temporal consistency in interactive video generation, which synthesizes dynamic and controllable video worlds th…