4 citations · 4 across the 1 of their papers we have counts for
2 papers
eess.AS2021★ 4 cited
TitaNet: Neural Model for speaker representation with 1D Depth-wise separable convolutions and global context
Nithin Rao Koluguri, Taejin Park, Boris Ginsburg
In this paper, we propose TitaNet, a novel neural network architecture for extracting speaker representations. We employ 1D depth-wise separable convolutions with Squeeze-and-Excit…
cs.SD2020
Robust Multi-channel Speech Recognition using Frequency Aligned Network
Taejin Park, Kenichi Kumatani, Minhua Wu +1
Conventional speech enhancement technique such as beamforming has known benefits for far-field speech recognition. Our own work in frequency-domain multi-channel acoustic modeling…