most citedExploring Speech Recognition, Translation, and Understanding with Discrete Speech Units: A Comparative Study

2 citations · 2 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CL2023

Reproducing Whisper-Style Training Using an Open-Source Toolkit and Publicly Available Data

Yifan Peng, Jinchuan Tian, Brian Yan +13

Pre-training speech models on large volumes of data has achieved remarkable success. OpenAI Whisper is a multilingual multitask model trained on 680k hours of supervised speech dat…

cs.CL2023

Evaluating Speech Synthesis by Training Recognizers on Synthetic Speech

Dareen Alharthi, Roshan Sharma, Hira Dhamyal +3

Modern speech synthesis systems have improved significantly, with synthetic speech being indistinguishable from real speech. However, efficient and holistic evaluation of synthetic…

cs.CL20232 cited

Exploring Speech Recognition, Translation, and Understanding with Discrete Speech Units: A Comparative Study

Xuankai Chang, Brian Yan, Kwanghee Choi +14

Speech signals, typically sampled at rates in the tens of thousands per second, contain redundancies, evoking inefficiencies in sequence modeling. High-dimensional speech features…

cs.CL2023

BASS: Block-wise Adaptation for Speech Summarization

Roshan Sharma, Kenneth Zheng, Siddhant Arora +3

End-to-end speech summarization has been shown to improve performance over cascade baselines. However, such models are difficult to train on very large inputs (dozens of minutes or…

cs.PL2023

Proving Logical Atomicity using Lock Invariants

Roshan Sharma, Shengyi Wang, Alexander Oey +3

Logical atomicity has been widely accepted as a specification format for data structures in concurrent separation logic. While both lock-free and lock-based data structures have be…