activity
20242026
collaborators

8 papers

cs.CV2026

EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal

Feier Wu, Wanke Xia, Xu He +8

Video object removal must eliminate not only the target object but also its induced effects while maintaining high-fidelity and spatiotemporally coherent restoration. Existing meth…

cs.SD2026

The Interspeech 2026 Audio Encoder Capability Challenge for Large Audio Language Models

Heinrich Dinkel, Jiahao Zhou, Guanbo Wang +8

This paper presents the Interspeech 2026 Audio Encoder Capability Challenge, a benchmark specifically designed to evaluate and advance the performance of pre-trained audio encoders…

cs.CV2026

PROMO: Promptable Outfitting for Efficient High-Fidelity Virtual Try-On

Haohua Chen, Tianze Zhou, Wei Zhu +8

Virtual Try-on (VTON) has become a core capability for online retail, where realistic try-on results provide reliable fit guidance, reduce returns, and benefit both consumers and m…

cs.CV2026

Kling-MotionControl Technical Report

Kling Team, Jialu Chen, Yikang Ding +21

Character animation aims to generate lifelike videos by transferring motion dynamics from a driving video to a reference image. Recent strides in generative models have paved the w…

cs.CV2025

From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing via Generative Bootstrapping

Xu He, Haoxian Zhang, Hejia Chen +7

Audio-driven visual dubbing aims to synchronize a video's lip movements with new speech but is fundamentally challenged by the lack of ideal training data: paired videos differing…

cs.CV2025

StableDub: Taming Diffusion Prior for Generalized and Efficient Visual Dubbing

Liyang Chen, Tianze Zhou, Xu He +7

The visual dubbing task aims to generate mouth movements synchronized with the driving audio, which has seen significant progress in recent years. However, two critical deficiencie…