1 citations · 2 across the 3 of their papers we have counts for
4 papers
TADS: Task-Aware Data Selection for Multi-Task Multimodal Pre-Training
Guanjie Cheng, Boyi Li, Lingyu Sun +4
Large-scale multimodal pre-trained models like CLIP rely heavily on high-quality training data, yet raw web-crawled datasets are often noisy, misaligned, and redundant, leading to…
Proficiency-Aware Adaptation and Data Augmentation for Robust L2 ASR
Ling Sun, Charlotte Zhu, Shuju Shi
General-purpose ASR underperforms for atypical speakers, such as L2 learners, reinforcing bias and limiting use in education and accessibility. Using the CEFR-graded Speak and Impr…
Efficient Emotional Adaptation for Audio-Driven Talking-Head Generation
Yuan Gan, Zongxin Yang, Xihang Yue +2
Audio-driven talking-head synthesis is a popular research topic for virtual human-related applications. However, the inflexibility and inefficiency of existing methods, which neces…
MelodyGLM: Multi-task Pre-training for Symbolic Melody Generation
Xinda Wu, Zhijie Huang, Kejun Zhang +5
Pre-trained language models have achieved impressive results in various music understanding and generation tasks. However, existing pre-training methods for symbolic melody generat…