8 papers
Using LLMs for Late Multimodal Sensor Fusion for Activity Recognition
Ilker Demirel, Karan Thakkar, Benjamin Elizalde +7
Sensor data streams provide valuable information around activities and context for downstream applications, though integrating complementary information can be challenging. We show…
Speech Foundation Models Generalize to Time Series Tasks from Wearable Sensor Data
Jaya Narain, Zakaria Aldeneh, Shirley Ren
Both speech and sensor time series data encode information in both the time- and frequency- domains, like spectral powers and waveform shapelets. We show that speech foundation mod…
Switchboard-Affect: Emotion Perception Labels from Conversational Speech
Amrit Romana, Jaya Narain, Tien Dung Tran +4
Understanding the nuances of speech emotion dataset curation and labeling is essential for assessing speech emotion recognition (SER) model potential in real-world applications. Mo…
Affect Models Have Weak Generalizability to Atypical Speech
Jaya Narain, Amrit Romana, Vikramjit Mitra +2
Speech and voice conditions can alter the acoustic properties of speech, which could impact the performance of paralinguistic models for affect for people with atypical speech. We…
Voice Quality Dimensions as Interpretable Primitives for Speaking Style for Atypical Speech and Affect
Jaya Narain, Vasudha Kowtha, Colin Lea +8
Perceptual voice quality dimensions describe key characteristics of atypical speech and other speech modulations. Here we develop and evaluate voice quality models for seven voice…
RelCon: Relative Contrastive Learning for a Motion Foundation Model for Wearable Data
Maxwell A. Xu, Jaya Narain, Gregory Darnell +7
We present RelCon, a novel self-supervised Relative Contrastive learning approach for training a motion foundation model from wearable accelerometry sensors. First, a learnable dis…