collaborators

11 papers

cs.CL2026

Quantifying Speaker Embedding Phonological Rule Interactions in Accented Speech Synthesis

Thanathai Lertpetchpun, Yoonjeong Lee, Thanapat Trachu +4

Many spoken languages, including English, exhibit wide variation in dialects and accents, making accent control an important capability for flexible text-to-speech (TTS) models. Cu…

cs.LG2025

Neural Codecs as Biosignal Tokenizers

Kleanthis Avramidis, Tiantian Feng, Woojae Jeong +4

Neurophysiological recordings such as electroencephalography (EEG) offer accessible and minimally invasive means of estimating physiological activity for applications in healthcare…

cs.SD2025

A long-form single-speaker real-time MRI speech dataset and benchmark

Sean Foley, Jihwan Lee, Kevin Huang +4

We release the USC Long Single-Speaker (LSS) dataset containing real-time MRI video of the vocal tract dynamics and simultaneous audio obtained during speech production. This uniqu…

eess.AS2025

ARTI-6: Towards Six-dimensional Articulatory Speech Encoding

Jihwan Lee, Sean Foley, Thanathai Lertpetchpun +6

We propose ARTI-6, a compact six-dimensional articulatory speech encoding framework derived from real-time MRI data that captures crucial vocal tract regions including the velum, t…

eess.IV2025

Interpretable Modeling of Articulatory Temporal Dynamics from real-time MRI for Phoneme Recognition

Jay Park, Hong Nguyen, Sean Foley +4

Real-time Magnetic Resonance Imaging (rtMRI) visualizes vocal tract action, offering a comprehensive window into speech articulation. However, its signals are high dimensional and…

cs.SD2025

Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe

Tiantian Feng, Kevin Huang, Anfeng Xu +6

We present Voxlect, a novel benchmark for modeling dialects and regional languages worldwide using speech foundation models. Specifically, we report comprehensive benchmark evaluat…