1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.CL2023
Reproducing Whisper-Style Training Using an Open-Source Toolkit and Publicly Available Data
Yifan Peng, Jinchuan Tian, Brian Yan +13
Pre-training speech models on large volumes of data has achieved remarkable success. OpenAI Whisper is a multilingual multitask model trained on 680k hours of supervised speech dat…
cs.SD2023
Retraining-free Customized ASR for Enharmonic Words Based on a Named-Entity-Aware Model and Phoneme Similarity Estimation
Yui Sudo, Kazuya Hata, Kazuhiro Nakadai
End-to-end automatic speech recognition (E2E-ASR) has the potential to improve performance, but a specific issue that needs to be addressed is the difficulty it has in handling enh…
cs.CL2023★ 1 cited
DPHuBERT: Joint Distillation and Pruning of Self-Supervised Speech Models
Yifan Peng, Yui Sudo, Shakeel Muhammad +1
Self-supervised learning (SSL) has achieved notable success in many speech processing tasks, but the large model size and heavy computational cost hinder the deployment. Knowledge…