activity
20242026
collaborators

8 papers

eess.AS2026

Adaptive Federated Fine-Tuning of Self-Supervised Speech Representations

Xin Guo, Chunrui Zhao, Hong Jia +4

Integrating Federated Learning (FL) with self-supervised learning (SSL) enables privacy-preserving fine-tuning for speech tasks. However, federated environments exhibit significant…

cs.CV2026

Semantic Audio-Visual Navigation in Continuous Environments

Yichen Zeng, Hebaixu Wang, Meng Liu +4

Audio-visual navigation enables embodied agents to navigate toward sound-emitting targets by leveraging both auditory and visual cues. However, most existing approaches rely on pre…

cs.SD2026

Edge-Cloud Collaborative Speech Emotion Captioning via Token-Level Speculative Decoding in Audio-Language Models

Xiangyuan Xue, Jiajun Lu, Yan Gao +3

Speech Emotion Captioning (SEC) leverages large audio-language models to generate rich, context-aware affective descriptions from speech. However, real-world deployment remains cha…

eess.AS2026

DTT-BSR: GAN-based DTTNet with RoPE Transformer Enhancement for Music Source Restoration

Shihong Tan, Haoyu Wang, Youran Ni +8

Music source restoration (MSR) aims to recover unprocessed stems from mixed and mastered recordings. The challenge lies in both separating overlapping sources and reconstructing si…

cs.SD2025

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation

Xiaoran Yang, Jianxuan Yang, Xinyue Guo +3

A key challenge in synthesizing audios from silent videos is the inherent trade-off between synthesis quality and inference efficiency in existing methods. For instance, flow match…

cs.MM2025

MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization

Jianxuan Yang, Xiaoran Yang, Lipan Zhang +3

Current video-to-audio (V2A) methods struggle in complex multi-event scenarios (video scenarios involving multiple sound sources, sound events, or transitions) due to two critical…