collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

Hierarchical Dual-Subspace Decoupling for Continual Learning in Vision-Language Models

Mengxin Qin, Xiang Zhang, Kun Wei +2

Class-incremental learning aims to continuously acquire new knowledge while preserving previously learned information, thereby mitigating catastrophic forgetting. Existing methods…

cs.CV2026

DIMoE-Adapters: Dynamic Expert Evolution for Continual Learning in Vision-Language Models

Mengxin Qin, Xiang Zhang, Xi Wang +3

Continual learning enables vision-language models to accumulate knowledge and adapt to evolving tasks without retraining from scratch. However, in multi-domain task-incremental lea…

cs.CV2025

A Bridge from Audio to Video: Phoneme-Viseme Alignment Allows Every Face to Speak Multiple Languages

Zibo Su, Kun Wei, Jiahua Li +3

Speech-driven talking face synthesis (TFS) focuses on generating lifelike facial animations from speech input. Current TFS models perform well in English but struggle with non-Engl…

cs.CV2025

Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents

Jiahua Li, Zhanhe Zhang, Chenghao Xu +4

Long videos, characterized by temporal complexity and sparse task-relevant information, pose significant reasoning challenges for AI systems. Although existing Large Language Model…

cs.CV20241 cited

Do You Guys Want to Dance: Zero-Shot Compositional Human Dance Generation with Multiple Persons

Zhe Xu, Kun Wei, Xu Yang +1

Human dance generation (HDG) aims to synthesize realistic videos from images and sequences of driving poses. Despite great success, existing methods are limited to generating video…