17 citations · 36 across the 10 of their papers we have counts for
9 papers · 1 filter
Improving Practical Aspects of End-to-End Multi-Talker Speech Recognition for Online and Offline Scenarios
Aswin Shanmugam Subramanian, Amit Das, Naoyuki Kanda +3
We extend the frameworks of Serialized Output Training (SOT) to address practical needs of both streaming and offline automatic speech recognition (ASR) applications. Our approach…
Hyperbolic Audio Source Separation
Darius Petermann, Gordon Wichern, Aswin Subramanian +1
We introduce a framework for audio source separation using embeddings on a hyperbolic manifold that compactly represent the hierarchical relationship between sound sources and time…
Reverberation as Supervision for Speech Separation
Rohith Aralikatti, Christoph Boeddeker, Gordon Wichern +2
This paper proposes reverberation as supervision (RAS), a novel unsupervised loss function for single-channel reverberant speech separation. Prior methods for unsupervised separati…
The 2020 ESPnet update: new features, broadened applications, performance improvements, and future plans
Shinji Watanabe, Florian Boyer, Xuankai Chang +12
This paper describes the recent development of ESPnet (https://github.com/espnet/espnet), an end-to-end speech processing toolkit. This project was initiated in December 2017 to ma…
Directional ASR: A New Paradigm for E2E Multi-Speaker Speech Recognition with Source Localization
Aswin Shanmugam Subramanian, Chao Weng, Shinji Watanabe +4
This paper proposes a new paradigm for handling far-field multi-speaker data in an end-to-end neural network manner, called directional automatic speech recognition (D-ASR), which…
The JHU Multi-Microphone Multi-Speaker ASR System for the CHiME-6 Challenge
Ashish Arora, Desh Raj, Aswin Shanmugam Subramanian +7
This paper summarizes the JHU team's efforts in tracks 1 and 2 of the CHiME-6 challenge for distant multi-microphone conversational speech diarization and recognition in everyday h…