3 citations · 6 across the 5 of their papers we have counts for
4 papers · 1 filter
Reproducing Whisper-Style Training Using an Open-Source Toolkit and Publicly Available Data
Yifan Peng, Jinchuan Tian, Brian Yan +13
Pre-training speech models on large volumes of data has achieved remarkable success. OpenAI Whisper is a multilingual multitask model trained on 680k hours of supervised speech dat…
Exploring Speech Recognition, Translation, and Understanding with Discrete Speech Units: A Comparative Study
Xuankai Chang, Brian Yan, Kwanghee Choi +14
Speech signals, typically sampled at rates in the tens of thousands per second, contain redundancies, evoking inefficiencies in sequence modeling. High-dimensional speech features…
BASS: Block-wise Adaptation for Speech Summarization
Roshan Sharma, Kenneth Zheng, Siddhant Arora +3
End-to-end speech summarization has been shown to improve performance over cascade baselines. However, such models are difficult to train on very large inputs (dozens of minutes or…
A Summary of the First Workshop on Language Technology for Language Documentation and Revitalization
Graham Neubig, Shruti Rijhwani, Alexis Palmer +21
Despite recent advances in natural language processing and other language technology, the application of such technology to language documentation and conservation has been limited…