4 papers
ARTI-6: Towards Six-dimensional Articulatory Speech Encoding
Jihwan Lee, Sean Foley, Thanathai Lertpetchpun +6
We propose ARTI-6, a compact six-dimensional articulatory speech encoding framework derived from real-time MRI data that captures crucial vocal tract regions including the velum, t…
A long-form single-speaker real-time MRI speech dataset and benchmark
Sean Foley, Jihwan Lee, Kevin Huang +4
We release the USC Long Single-Speaker (LSS) dataset containing real-time MRI video of the vocal tract dynamics and simultaneous audio obtained during speech production. This uniqu…
Towards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition
Sean Foley, Hong Nguyen, Jihwan Lee +4
Although many previous studies have carried out multimodal learning with real-time MRI data that captures the audio-visual kinematics of the vocal tract during speech, these studie…
Articulatory Feature Prediction from Surface EMG during Speech Production
Jihwan Lee, Kevin Huang, Kleanthis Avramidis +6
We present a model for predicting articulatory features from surface electromyography (EMG) signals during speech production. The proposed model integrates convolutional layers and…