collaborators

6 papers

eess.AS2026

Autoencoder based optimized SSL representations: Complexity Minimization and improved Dysarthric ASR

Paban Sapkota, Hemant Kumar Kathania, Mikko Kurimo +2

Self-supervised learning (SSL) models extract rich speech representations but often come with high-dimensional features, increasing computational complexity. This work explores an…

eess.AS2026

ARTI-6: Towards Six-dimensional Articulatory Speech Encoding

Jihwan Lee, Sean Foley, Thanathai Lertpetchpun +6

We propose ARTI-6, a compact six-dimensional articulatory speech encoding framework derived from real-time MRI data that captures crucial vocal tract regions including the velum, t…

cs.SD2025

A long-form single-speaker real-time MRI speech dataset and benchmark

Sean Foley, Jihwan Lee, Kevin Huang +4

We release the USC Long Single-Speaker (LSS) dataset containing real-time MRI video of the vocal tract dynamics and simultaneous audio obtained during speech production. This uniqu…

eess.AS2025

On the Relationship between Accent Strength and Articulatory Features

Kevin Huang, Sean Foley, Jihwan Lee +3

This paper explores the relationship between accent strength and articulatory features inferred from acoustic speech. To quantify accent strength, we compare phonetic transcription…

eess.AS2025

Articulatory Feature Prediction from Surface EMG during Speech Production

Jihwan Lee, Kevin Huang, Kleanthis Avramidis +6

We present a model for predicting articulatory features from surface electromyography (EMG) signals during speech production. The proposed model integrates convolutional layers and…

cs.SD2025

Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits

Tiantian Feng, Jihwan Lee, Anfeng Xu +9

We introduce Vox-Profile, a comprehensive benchmark to characterize rich speaker and speech traits using speech foundation models. Unlike existing works that focus on a single dime…