8 papers
Adaptive Federated Fine-Tuning of Self-Supervised Speech Representations
Xin Guo, Chunrui Zhao, Hong Jia +4
Integrating Federated Learning (FL) with self-supervised learning (SSL) enables privacy-preserving fine-tuning for speech tasks. However, federated environments exhibit significant…
Semantic Audio-Visual Navigation in Continuous Environments
Yichen Zeng, Hebaixu Wang, Meng Liu +4
Audio-visual navigation enables embodied agents to navigate toward sound-emitting targets by leveraging both auditory and visual cues. However, most existing approaches rely on pre…
Edge-Cloud Collaborative Speech Emotion Captioning via Token-Level Speculative Decoding in Audio-Language Models
Xiangyuan Xue, Jiajun Lu, Yan Gao +3
Speech Emotion Captioning (SEC) leverages large audio-language models to generate rich, context-aware affective descriptions from speech. However, real-world deployment remains cha…
DTT-BSR: GAN-based DTTNet with RoPE Transformer Enhancement for Music Source Restoration
Shihong Tan, Haoyu Wang, Youran Ni +8
Music source restoration (MSR) aims to recover unprocessed stems from mixed and mastered recordings. The challenge lies in both separating overlapping sources and reconstructing si…
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation
Xiaoran Yang, Jianxuan Yang, Xinyue Guo +3
A key challenge in synthesizing audios from silent videos is the inherent trade-off between synthesis quality and inference efficiency in existing methods. For instance, flow match…
MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization
Jianxuan Yang, Xiaoran Yang, Lipan Zhang +3
Current video-to-audio (V2A) methods struggle in complex multi-event scenarios (video scenarios involving multiple sound sources, sound events, or transitions) due to two critical…