4 papers
From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning
Kele Xu, Yulu Fang, Boda Zhou +6
This paper examines audio self-supervised learning (SSL) through the alignment between pretraining objectives, architectural inductive biases, and downstream applications. Rather t…
Timeripple: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space
Wenxuan Miao, Yulin Sun, Aiyue Chen +8
The recent surge in video generation has shown the growing demand for high-quality video synthesis using large vision models. Existing video generation models are predominantly bas…
AudioSet-R: A Refined AudioSet with Multi-Stage LLM Label Reannotation
Yulin Sun, Qisheng Xu, Yi Su +4
AudioSet is a widely used benchmark in the audio research community and has significantly advanced various audio-related tasks. However, persistent issues with label accuracy and c…
AudioCIL: A Python Toolbox for Audio Class-Incremental Learning with Multiple Scenes
Qisheng Xu, Yulin Sun, Yi Su +7
Deep learning, with its robust aotomatic feature extraction capabilities, has demonstrated significant success in audio signal processing. Typically, these methods rely on static,…