17 citations · 36 across the 7 of their papers we have counts for
10 papers
Hyperbolic Audio Source Separation
Darius Petermann, Gordon Wichern, Aswin Subramanian +1
We introduce a framework for audio source separation using embeddings on a hyperbolic manifold that compactly represent the hierarchical relationship between sound sources and time…
Reverberation as Supervision for Speech Separation
Rohith Aralikatti, Christoph Boeddeker, Gordon Wichern +2
This paper proposes reverberation as supervision (RAS), a novel unsupervised loss function for single-channel reverberant speech separation. Prior methods for unsupervised separati…
An Exploration of Self-Supervised Pretrained Representations for End-to-End Speech Recognition
Xuankai Chang, Takashi Maekaku, Pengcheng Guo +8
Self-supervised pretraining on speech data has achieved a lot of progress. High-fidelity representation of the speech signal is learned from a lot of untranscribed data and shows p…
The 2020 ESPnet update: new features, broadened applications, performance improvements, and future plans
Shinji Watanabe, Florian Boyer, Xuankai Chang +12
This paper describes the recent development of ESPnet (https://github.com/espnet/espnet), an end-to-end speech processing toolkit. This project was initiated in December 2017 to ma…
Directional ASR: A New Paradigm for E2E Multi-Speaker Speech Recognition with Source Localization
Aswin Shanmugam Subramanian, Chao Weng, Shinji Watanabe +4
This paper proposes a new paradigm for handling far-field multi-speaker data in an end-to-end neural network manner, called directional automatic speech recognition (D-ASR), which…
The JHU Multi-Microphone Multi-Speaker ASR System for the CHiME-6 Challenge
Ashish Arora, Desh Raj, Aswin Shanmugam Subramanian +7
This paper summarizes the JHU team's efforts in tracks 1 and 2 of the CHiME-6 challenge for distant multi-microphone conversational speech diarization and recognition in everyday h…