works on

From the 1 of 9 linked papers with an AI index.

collaborators

9 papers

cs.SD2026

OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder

Shikhar Bharadwaj, Samuele Cornell, Kwanghee Choi +4

OpenBEATs is an open-source framework that extends the BEATs audio encoder with multi-domain masked-token pretraining, achieving state-of-the-art results on a wide range of audio t…

eess.AS2026

ANCHOR: Autoregressive Non-intrusive Chunk-Ordered Refinement for Joint Multi-Resolution Speech Quality Modeling

Zhuoyan Tao, Jiatong Shi, Hye-jin Shim +1

While speech quality is typically assessed on complete utterances, streaming and generative systems require incremental estimation from partial audio. Existing predictors assume fu…

cs.SD2026

The CMU-AIST submission for the ICME 2025 Audio Encoder Challenge

Shikhar Bharadwaj, Samuele Cornell, Kwanghee Choi +4

This technical report describes our submission to the ICME 2025 audio encoder challenge. Our submitted system is built on BEATs, a masked speech token prediction based audio encode…

cs.SD2025

ARECHO: Autoregressive Evaluation via Chain-Based Hypothesis Optimization for Speech Multi-Metric Estimation

Jiatong Shi, Yifan Cheng, Bo-Hao Su +6

Speech signal analysis poses significant challenges, particularly in tasks such as speech quality evaluation and profiling, where the goal is to predict multiple perceptual and obj…

cs.CL2025

Geolocation-Aware Robust Spoken Language Identification

Qingzheng Wang, Hye-jin Shim, Jiancheng Sun +1

While Self-supervised Learning (SSL) has significantly improved Spoken Language Identification (LID), existing models often struggle to consistently classify dialects and accents o…

cs.LG2025

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet

Anyu Ying, Natarajan Balaji Shankar, Chyi-Jiunn Lin +7

Despite advancements in ASR, child speech recognition remains challenging due to acoustic variability and limited annotated data. While fine-tuning adult ASR models on child speech…