7 citations · 18 across the 6 of their papers we have counts for
6 papers
Large-scale learning of generalised representations for speaker recognition
Jee-weon Jung, Hee-Soo Heo, Bong-Jin Lee +5
The objective of this work is to develop a speaker recognition model to be used in diverse scenarios. We hypothesise that two components should be adequately configured to build su…
Better Intermediates Improve CTC Inference
Tatsuya Komatsu, Yusuke Fujita, Jaesong Lee +3
This paper proposes a method for improved CTC inference with searched intermediates and multi-pass conditioning. The paper first formulates self-conditioned CTC as a probabilistic…
Memory-Efficient Training of RNN-Transducer with Sampled Softmax
Jaesong Lee, Lukas Lee, Shinji Watanabe
RNN-Transducer has been one of promising architectures for end-to-end automatic speech recognition. Although RNN-Transducer has many advantages including its strong accuracy and st…
A Comparative Study on Non-Autoregressive Modelings for Speech-to-Text Generation
Yosuke Higuchi, Nanxin Chen, Yuya Fujita +6
Non-autoregressive (NAR) models simultaneously generate multiple outputs in a sequence, which significantly reduces the inference speed at the cost of accuracy drop compared to aut…
Layer Pruning on Demand with Intermediate CTC
Jaesong Lee, Jingu Kang, Shinji Watanabe
Deploying an end-to-end automatic speech recognition (ASR) model on mobile/embedded devices is a challenging task, since the device computational power and energy consumption requi…
Intermediate Loss Regularization for CTC-based Speech Recognition
Jaesong Lee, Shinji Watanabe
We present a simple and efficient auxiliary loss function for automatic speech recognition (ASR) based on the connectionist temporal classification (CTC) objective. The proposed ob…