4 papers
Differentiable Acoustic Radiance Transfer
Sungho Lee, Matteo Scerbo, Seungu Han +3
Geometric acoustics is an efficient framework for room acoustics modeling, governed by the canonical time-dependent rendering equation. Acoustic radiance transfer (ART) solves the…
Cross-Modal Bottleneck Fusion For Noise Robust Audio-Visual Speech Recognition
Seaone Ok, Min Jun Choi, Eungbeom Kim +2
Audio-Visual Speech Recognition (AVSR) leverages both acoustic and visual cues to improve speech recognition under noisy conditions. A central question is how to design a fusion me…
Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective
Seungu Han, Sungho Lee, Kyogu Lee
Recent speech enhancement (SE) models increasingly leverage self-supervised learning (SSL) representations for their rich semantic information. Typically, intermediate features are…
Few-step Adversarial Schrödinger Bridge for Generative Speech Enhancement
Seungu Han, Sungho Lee, Juheon Lee +1
Deep generative models have recently been employed for speech enhancement to generate perceptually valid clean speech on large-scale datasets. Several diffusion models have been pr…