11 papers
Ambisonics Encoding of Room Impulse Responses using a Device-Agnostic Diffusion Model
Eloi Moliner, Christoph Hold, Juan Azcarreta Ortiz +5
We address the problem of encoding room impulse responses (RIRs) into high-order Ambisonics (HOA) representations from arbitrary and potentially insufficient or incomplete micropho…
Conditional Flow Matching for Visually-Guided Acoustic Highlighting
Hugo Malard, Gael Le Lan, Daniel Wong +3
Visually-guided acoustic highlighting seeks to rebalance audio in alignment with the accompanying video, creating a coherent audio-visual experience. While visual saliency and enha…
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models
Hugo Malard, Michel Olvera, Sanjeel Parekh +3
Weak supervision sets a practical regime for audio-visual sound source localization as dense spatial annotations are costly to obtain at scale. The task, however, remains challengi…
Spatial-Magnifier: Spatial upsampling for multichannel speech enhancement
Dongheon Lee, Ashutosh Pandey, Sanjeel Parekh +4
While the spatial directivity of multichannel speech enhancement algorithms improves with the number of microphones, fitting large capture arrays into real-world edge devices is ty…
Sound Event Detection with Boundary-Aware Optimization and Inference
Florian Schmid, Chi Ian Tang, Sanjeel Parekh +9
Temporal detection problems appear in many fields including time-series estimation, activity recognition and sound event detection (SED). In this work, we propose a new approach to…
Unified Diffusion Refinement for Multi-Channel Speech Enhancement and Separation
Zhongweiyang Xu, Ashutosh Pandey, Juan Azcarreta +4
We propose Uni-ArrayDPS, a novel diffusion-based refinement framework for unified multi-channel speech enhancement and separation. Existing methods for multi-channel speech enhance…