most citedHunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

1 citations · 1 across the 9 of their papers we have counts for

collaborators

10 papers

cs.CV2026

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment

Zixiang Zhou, Zhentao Yu, Yifeng Ma +9

Subject-driven and multi-element video generation are central to controllable video synthesis, but existing methods still struggle to preserve identity consistency and model comple…

cs.CV2026

Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing

Sen Liang, Cong Wang, Zhentao Yu +8

Existing instruction-based video editing datasets commonly focus on single-task appearance editing, failing to meet the complex creative demands of real-world scenarios. To bridge…

cs.CV2026

HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation

Cong Wang, Zhentao Yu, Hongmei Wang +7

Current identity-consistent video generation methods struggle to preserve appearance fidelity under large viewpoint changes. While introducing multi-view reference input offers a n…

cs.CV2026

SpongeBob: Sync-Aware Harmonious Audio-Visual Generative Editing

Sen Liang, Cong Wang, Fengbin Guan +6

Visual and acoustic events in the physical world are inherently coupled, yet existing video editing methods typically adopt decoupled pipelines, lacking bidirectional modality inte…

cs.CV2026

Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking Avatars

Youliang Zhang, Zhengguang Zhou, Zhentao Yu +11

Generating talking avatars is a fundamental task in video generation. Although existing methods can generate full-body talking avatars with simple human motion, extending this task…

cs.CV2025

Harmony: Harmonizing Audio and Video Generation through Cross-Task Synergy

Teng Hu, Zhentao Yu, Guozhen Zhang +6

The synthesis of synchronized audio-visual content is a key challenge in generative AI, with open-source models facing challenges in robust audio-video alignment. Our analysis reve…