output
20122024
most citedCaptum: A unified and generic model interpretability library for PyTorch

649 citations

Showing cs.SDShow all

7 papers · 1 filter

cs.SD2022

The impact of removing head movements on audio-visual speech enhancement

Zhiqi Kang, Mostafa Sadeghi, Radu Horaud +3

This paper investigates the impact of head movements on audio-visual speech enhancement (AVSE). Although being a common conversational feature, head movements have been ignored by…

cs.SD20211 cited

Rendering Spatial Sound for Interoperable Experiences in the Audio Metaverse

Jean-Marc Jot, Rémi Audfray, Mark Hertensteiner +1

Interactive audio spatialization technology previously developed for video game authoring and rendering has evolved into an essential component of platforms enabling shared immersi…

cs.SD202132 cited

EasyCom: An Augmented Reality Dataset to Support Algorithms for Easy Communication in Noisy Environments

Jacob Donley, Vladimir Tourbabin, Jung-Suk Lee +6

Augmented Reality (AR) as a platform has the potential to facilitate the reduction of the cocktail party effect. Future AR headsets could potentially leverage information from an a…

cs.SD20211 cited

Flexi-Transducer: Optimizing Latency, Accuracy and Compute forMulti-Domain On-Device Scenarios

Jay Mahadeokar, Yangyang Shi, Yuan Shangguan +7

Often, the storage and computational constraints of embeddeddevices demand that a single on-device ASR model serve multiple use-cases / domains. In this paper, we propose aFlexible…

cs.SD20212 cited

Multi-rate attention architecture for fast streamable Text-to-speech spectrum modeling

Qing He, Zhiping Xiu, Thilo Koehler +1

Typical high quality text-to-speech (TTS) systems today use a two-stage architecture, with a spectrum model stage that generates spectral frames and a vocoder stage that generates…

cs.SD202020 cited

A Sequential Self Teaching Approach for Improving Generalization in Sound Event Recognition

Anurag Kumar, Vamsi Krishna Ithapu

An important problem in machine auditory perception is to recognize and detect sound events. In this paper, we propose a sequential self-teaching approach to learning sounds. Our m…