From the 2 of 10 linked papers with an AI index.
10 papers
An Empirical Recipe for Universal Phone Recognition
Shikhar Bharadwaj, Chin-Jou Li, Kwanghee Choi +4
The paper introduces PhoneticXEUS, a multilingual phone recognition system trained on large-scale data that achieves state-of-the-art error rates on both many languages and accente…
PRiSM: Benchmarking Phone Realization in Speech Models
Shikhar Bharadwaj, Chin-Jou Li, Yoonjae Kim +13
Phone recognition (PR) serves as the atomic interface for language-agnostic modeling for cross-lingual speech processing and phonetic analysis. Despite prolonged efforts in develop…
OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder
Shikhar Bharadwaj, Samuele Cornell, Kwanghee Choi +4
OpenBEATs is an open-source framework that extends the BEATs audio encoder with multi-domain masked-token pretraining, achieving state-of-the-art results on a wide range of audio t…
Phone Segmentation and Recognition through Phonological Activation Mapping
Shikhar Bharadwaj, Kwanghee Choi, Stephen McIntosh +8
Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separately. We argue that phonetic structure is already latent in the re…
The CMU-AIST submission for the ICME 2025 Audio Encoder Challenge
Shikhar Bharadwaj, Samuele Cornell, Kwanghee Choi +4
This technical report describes our submission to the ICME 2025 audio encoder challenge. Our submitted system is built on BEATs, a masked speech token prediction based audio encode…
POWSM: A Phonetic Open Whisper-Style Speech Foundation Model
Chin-Jou Li, Kalvin Chang, Shikhar Bharadwaj +5
Recent advances in spoken language processing have led to substantial progress in phonetic tasks such as automatic speech recognition (ASR), phone recognition (PR), grapheme-to-pho…