2 citations · 2 across the 5 of their papers we have counts for
5 papers
Reproducing Whisper-Style Training Using an Open-Source Toolkit and Publicly Available Data
Yifan Peng, Jinchuan Tian, Brian Yan +13
Pre-training speech models on large volumes of data has achieved remarkable success. OpenAI Whisper is a multilingual multitask model trained on 680k hours of supervised speech dat…
Evaluating Speech Synthesis by Training Recognizers on Synthetic Speech
Dareen Alharthi, Roshan Sharma, Hira Dhamyal +3
Modern speech synthesis systems have improved significantly, with synthetic speech being indistinguishable from real speech. However, efficient and holistic evaluation of synthetic…
Exploring Speech Recognition, Translation, and Understanding with Discrete Speech Units: A Comparative Study
Xuankai Chang, Brian Yan, Kwanghee Choi +14
Speech signals, typically sampled at rates in the tens of thousands per second, contain redundancies, evoking inefficiencies in sequence modeling. High-dimensional speech features…
BASS: Block-wise Adaptation for Speech Summarization
Roshan Sharma, Kenneth Zheng, Siddhant Arora +3
End-to-end speech summarization has been shown to improve performance over cascade baselines. However, such models are difficult to train on very large inputs (dozens of minutes or…
Proving Logical Atomicity using Lock Invariants
Roshan Sharma, Shengyi Wang, Alexander Oey +3
Logical atomicity has been widely accepted as a specification format for data structures in concurrent separation logic. While both lock-free and lock-based data structures have be…