3 citations · 6 across the 12 of their papers we have counts for
3 papers · 1 filter
Unlocking Foundation Models for Privacy-Enhancing Speech Understanding: An Early Study on Low Resource Speech Training Leveraging Label-guided Synthetic Speech Content
Tiantian Feng, Digbalay Bose, Xuan Shi +1
Automatic Speech Understanding (ASU) leverages the power of deep learning models for accurate interpretation of human speech, leading to a wide range of speech applications that en…
On the Role of Visual Context in Enriching Music Representations
Kleanthis Avramidis, Shanti Stewart, Shrikanth Narayanan
Human perception and experience of music is highly context-dependent. Contextual variability contributes to differences in how we interpret and interact with music, challenging the…
Acted vs. Improvised: Domain Adaptation for Elicitation Approaches in Audio-Visual Emotion Recognition
Haoqi Li, Yelin Kim, Cheng-Hao Kuo +1
Key challenges in developing generalized automatic emotion recognition systems include scarcity of labeled data and lack of gold-standard references. Even for the cues that are lab…