6 citations · 6 across the 2 of their papers we have counts for
4 papers · 1 filter
E-Branchformer: Branchformer with Enhanced merging for speech recognition
Kwangyoun Kim, Felix Wu, Yifan Peng +4
Conformer, combining convolution and self-attention sequentially to capture both local and global information, has shown remarkable performance and is currently regarded as the sta…
Multi-mode Transformer Transducer with Stochastic Future Context
Kwangyoun Kim, Felix Wu, Prashant Sridhar +2
Automatic speech recognition (ASR) models make fewer errors when more surrounding speech information is presented as context. Unfortunately, acquiring a larger future context leads…
Tuplemax Loss for Language Identification
Li Wan, Prashant Sridhar, Yang Yu +2
In many scenarios of a language identification task, the user will specify a small set of languages which he/she can speak instead of a large set of all possible languages. We want…
VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking
Quan Wang, Hannah Muckenhirn, Kevin Wilson +7
In this paper, we present a novel system that separates the voice of a target speaker from multi-speaker signals, by making use of a reference signal from the target speaker. We ac…