5 papers
Spectral-Target Physical Latent Structuring for JEPA-Style World Models
Penghao Zhu, Salvatore Penachio, Kaustav Mukherjee +1
Latent world models have become increasingly popular as a method to predict and plan in latent space rather than pixel space. Recent architectures, such as LeWorldModel (LeWM), joi…
Multi-Task Learning for Non-Canonical Phoneme Recognition via Articulatory Feature Decomposition
Sophia Riaz, Haoze Zheng, Amos Roche +5
Pathological and more broadly non-canonical speech present significant challenges for automatic phoneme recognition due to systematic deviations from canonical pronunciation and li…
Multimodal Hidden Markov Models for Persistent Emotional State Tracking
Anamika Ragu, Aneesh Jonelagadda
Tracking an interpretable emotional arc of a conversation via the sentiment of individual utterances processed as a whole is central to both understanding and guiding communication…
Uncertainty-Aware Multimodal Emotion Recognition through Dirichlet Parameterization
Rémi Grzeczkowicz, Eric Soriano, Ali Janati +4
In this work, we present a lightweight and privacy-preserving Multimodal Emotion Recognition (MER) framework designed for deployment on edge devices. To demonstrate framework's ver…
Mnemosyne: An Unsupervised, Human-Inspired Long-Term Memory Architecture for Edge-Based LLMs
Aneesh Jonelagadda, Christina Hahn, Haoze Zheng +1
Long-term memory is essential for natural, realistic dialogue. However, current large language model (LLM) memory systems rely on either brute-force context expansion or static ret…