collaborators

6 papers

eess.AS2026

From Hallucination to Articulation: Language Model-Driven Losses for Ultra Low-Bitrate Neural Speech Coding

Jayeon Yi, Minje Kim

``Phoneme Hallucinations (PH)'' commonly occur in low-bitrate DNN-based codecs. It is the generative decoder's attempt to synthesize plausible outputs from excessively compressed t…

cs.SD2025

PromptSep: Generative Audio Separation via Multimodal Prompting

Yutong Wen, Ke Chen, Prem Seetharaman +7

Recent breakthroughs in language-queried audio source separation (LASS) have shown that generative models can achieve higher separation audio quality than traditional masking-based…

cs.SD2025

Low-Resource Audio Codec (LRAC): 2025 Challenge Description

Kamil Wojcicki, Yusuf Ziya Isik, Laura Lechler +8

While recent neural audio codecs deliver superior speech quality at ultralow bitrates over traditional methods, their practical adoption is hindered by obstacles related to low-res…

cs.SD2025

Combolutional Neural Networks

Cameron Churchwell, Minje Kim, Paris Smaragdis

Selecting appropriate inductive biases is an essential step in the design of machine learning models, especially when working with audio, where even short clips may contain million…

eess.AS2025

Adaptive Slimming for Scalable and Efficient Speech Enhancement

Riccardo Miccini, Minje Kim, Clément Laroche +2

Speech enhancement (SE) enables robust speech recognition, real-time communication, hearing aids, and other applications where speech quality is crucial. However, deploying such sy…

cs.SD2025

User-guided Generative Source Separation

Yutong Wen, Minje Kim, Paris Smaragdis

Music source separation (MSS) aims to extract individual instrument sources from their mixture. While most existing methods focus on the widely adopted four-stem separation setup (…