5 papers
Lightweight and Generalizable Acoustic Scene Representations via Contrastive Fine-Tuning and Distillation
Kuang Yuan, Yang Gao, Xilin Li +4
Acoustic scene classification (ASC) models on edge devices typically operate under fixed class assumptions, lacking the transferability needed for real-world applications that requ…
MOVA: Towards Scalable and Synchronized Video-Audio Generation
OpenMOSS Team, Donghua Yu, Mingshu Chen +38
Audio is indispensable for real-world video, yet generation models have largely overlooked audio components. Current approaches to producing audio-visual content often rely on casc…
Learning Brain Representation with Hierarchical Visual Embeddings
Jiawen Zheng, Haonan Jia, Ming Li +4
Decoding visual representations from brain signals has attracted significant attention in both neuroscience and artificial intelligence. However, the degree to which brain signals…
Residual Learning for Neural Ambisonics Encoders
Thomas Deppisch, Yang Gao, Manan Mittal +4
Emerging wearable devices such as smartglasses and extended reality headsets demand high-quality spatial audio capture from compact, head-worn microphone arrays. Ambisonics provide…
ArtiFree: Detecting and Reducing Generative Artifacts in Diffusion-based Speech Enhancement
Bhawana Chhaglani, Yang Gao, Julius Richter +4
Diffusion-based speech enhancement (SE) achieves natural-sounding speech and strong generalization, yet suffers from key limitations like generative artifacts and high inference la…