40 citations · 51 across the 3 of their papers we have counts for
4 papers
SPGISpeech: 5,000 hours of transcribed financial audio for fully formatted end-to-end speech recognition
Patrick K. O'Neill, Vitaly Lavrukhin, Somshubra Majumdar +10
In the English speech-to-text (STT) machine learning task, acoustic models are conventionally trained on uncased Latin characters, and any necessary orthography (such as capitaliza…
Tiny Transducer: A Highly-efficient Speech Recognition Model on Edge Devices
Yuekai Zhang, Sining Sun, Long Ma
This paper proposes an extremely lightweight phone-based transducer model with a tiny decoding graph on edge devices. First, a phone synchronous decoding (PSD) algorithm based on b…
Recent Developments on ESPnet Toolkit Boosted by Conformer
Pengcheng Guo, Florian Boyer, Xuankai Chang +12
In this study, we present recent developments on ESPnet: End-to-End Speech Processing toolkit, which mainly involves a recently proposed architecture called Conformer, Convolution-…
Sequence-to-sequence Singing Voice Synthesis with Perceptual Entropy Loss
Jiatong Shi, Shuai Guo, Nan Huo +2
The neural network (NN) based singing voice synthesis (SVS) systems require sufficient data to train well and are prone to over-fitting due to data scarcity. However, we often enco…