649 citations
- Carnegie Mellon UniversityUS23 papers
- Stanford UniversityUS20 papers
- Google (United States)US14 papers
- Georgia Institute of TechnologyUS13 papers
- Tel Aviv UniversityIL12 papers
- Cornell UniversityUS11 papers
- University of California, BerkeleyUS11 papers
- University College LondonGB10 papers
- Harvard University PressUS9 papers
- Johns Hopkins UniversityUS9 papers
- Massachusetts Institute of TechnologyUS9 papers
- The University of Texas at AustinUS9 papers
7 papers · 1 filter
The impact of removing head movements on audio-visual speech enhancement
Zhiqi Kang, Mostafa Sadeghi, Radu Horaud +3
This paper investigates the impact of head movements on audio-visual speech enhancement (AVSE). Although being a common conversational feature, head movements have been ignored by…
Rendering Spatial Sound for Interoperable Experiences in the Audio Metaverse
Jean-Marc Jot, Rémi Audfray, Mark Hertensteiner +1
Interactive audio spatialization technology previously developed for video game authoring and rendering has evolved into an essential component of platforms enabling shared immersi…
EasyCom: An Augmented Reality Dataset to Support Algorithms for Easy Communication in Noisy Environments
Jacob Donley, Vladimir Tourbabin, Jung-Suk Lee +6
Augmented Reality (AR) as a platform has the potential to facilitate the reduction of the cocktail party effect. Future AR headsets could potentially leverage information from an a…
Flexi-Transducer: Optimizing Latency, Accuracy and Compute forMulti-Domain On-Device Scenarios
Jay Mahadeokar, Yangyang Shi, Yuan Shangguan +7
Often, the storage and computational constraints of embeddeddevices demand that a single on-device ASR model serve multiple use-cases / domains. In this paper, we propose aFlexible…
Multi-rate attention architecture for fast streamable Text-to-speech spectrum modeling
Qing He, Zhiping Xiu, Thilo Koehler +1
Typical high quality text-to-speech (TTS) systems today use a two-stage architecture, with a spectrum model stage that generates spectral frames and a vocoder stage that generates…
A Sequential Self Teaching Approach for Improving Generalization in Sound Event Recognition
Anurag Kumar, Vamsi Krishna Ithapu
An important problem in machine auditory perception is to recognize and detect sound events. In this paper, we propose a sequential self-teaching approach to learning sounds. Our m…