From the 2 of 23 linked papers with an AI index.
23 papers
Dissecting Sensitivity to Training Language in Self-Supervised Speech Learning Using Neural Audio Codec Tokens
Daigo Takizawa, Tomohiko Nakamura, Samuele Cornell +3
The paper investigates how the language used to train neural audio codecs and self‑supervised speech models affects performance, finding that codec training language has little imp…
OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder
Shikhar Bharadwaj, Samuele Cornell, Kwanghee Choi +4
OpenBEATs is an open-source framework that extends the BEATs audio encoder with multi-domain masked-token pretraining, achieving state-of-the-art results on a wide range of audio t…
ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era
Masao Someki, Alexander Polok, Carlos Carvalho +14
Recent speech research involves increasingly large datasets, complex models, and diverse experimental workflows. However, existing frameworks require substantial engineering effort…
Grounding Spoken LLMs in Multi-Speaker Audio via Diarization Conditioning
Alexander Polok, Samuele Cornell, Sathvik Udupa +3
We propose diarization-conditioned spoken language models (SLMs), a strategy for extending SLMs to far-field multi-talker audio. Rather than adapting the decoder via Serialized Out…
Exploiting Noise Inseparability for Weakly-Supervised Discriminative Speech Denoising Using Noisy Targets
Matthew Maciejewski, Samuele Cornell
Speech denoising is an often necessary step not only for human listening, but also for downstream processing by systems lacking robustness to noisy, real-world acoustic conditions.…
Cross-Talk Speech Reduction, by Separation, for Separation
Zhong-Qiu Wang, Samuele Cornell
In conversational speech separation and recognition tasks, close-talk microphones are typically attached to each speaker during training data collection to capture near-field, clos…