works on

From the 2 of 30 linked papers with an AI index.

activity
20242026
collaborators
Showing eess.ASShow all

20 papers · 1 filter

eess.AS2026

ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era

Masao Someki, Alexander Polok, Carlos Carvalho +14

Recent speech research involves increasingly large datasets, complex models, and diverse experimental workflows. However, existing frameworks require substantial engineering effort…

eess.AS2026

Grounding Spoken LLMs in Multi-Speaker Audio via Diarization Conditioning

Alexander Polok, Samuele Cornell, Sathvik Udupa +3

We propose diarization-conditioned spoken language models (SLMs), a strategy for extending SLMs to far-field multi-talker audio. Rather than adapting the decoder via Serialized Out…

eess.AS2026

Exploiting Noise Inseparability for Weakly-Supervised Discriminative Speech Denoising Using Noisy Targets

Matthew Maciejewski, Samuele Cornell

Speech denoising is an often necessary step not only for human listening, but also for downstream processing by systems lacking robustness to noisy, real-world acoustic conditions.…

eess.AS2026

Cross-Talk Speech Reduction, by Separation, for Separation

Zhong-Qiu Wang, Samuele Cornell

In conversational speech separation and recognition tasks, close-talk microphones are typically attached to each speaker during training data collection to capture near-field, clos…

eess.AS2026

Ring Mixing with Auxiliary Signal-to-Consistency-Error Ratio Loss for Unsupervised Denoising in Speech Separation

Matthew Maciejewski, Samuele Cornell

Noisy speech separation systems are typically trained on fully-synthetic mixtures, limiting generalization to real-world scenarios. Though training on mixtures of in-domain (thus o…

eess.AS2026

MAPSS: Manifold-based Assessment of Perceptual Source Separation

Amir Ivry, Samuele Cornell, Shinji Watanabe

Objective assessment of audio source-separation systems still mismatches subjective human perception, especially when interference from competing talkers and distortion of the targ…