14 papers
ULTRAS -- Unified Learning of Transformer Representations for Audio and Speech Signals
Ameenudeen P E, Charumathi Narayanan, Sriram Ganapathy
Self-supervised learning (SSL) has driven impressive advances in speech processing by adopting time-domain prediction objectives, while audio representation learning frameworks ope…
Benchmarking Humans and Machines on Complex Multilingual Speech Understanding Tasks
Sai Samrat Kankanala, Ram Chandra, Sriram Ganapathy
Auditory attention and selective phase-locking are central to human speech understanding in complex acoustic scenes and cocktail party settings, yet these capabilities in multiling…
Textless and Non-Parallel Speech-to-Speech Emotion Style Transfer
Soumya Dutta, Avni Jain, Sriram Ganapathy
Given a pair of source and reference speech recordings, speech-to-speech (S2S) emotion style transfer involves the generation of an output speech that mimics the emotion characteri…
Benchmarking Speech Systems for Frontline Health Conversations: The DISPLACE-M Challenge
Dhanya E, Ankita Meena, Manas Nanivadekar +11
The DIarization and Speech Processing for LAnguage understanding in Conversational Environments - Medical (DISPLACE-M) challenge introduces a conversational AI benchmark for unders…
A Mixture-of-Experts Model for Multimodal Emotion Recognition in Conversations
Soumya Dutta, Smruthi Balaji, Sriram Ganapathy
Emotion Recognition in Conversations (ERC) presents unique challenges, requiring models to capture the temporal flow of multi-turn dialogues and to effectively integrate cues from…
FESTA: Functionally Equivalent Sampling for Trust Assessment of Multimodal LLMs
Debarpan Bhattacharya, Apoorva Kulkarni, Sriram Ganapathy
The accurate trust assessment of multimodal large language models (MLLMs) generated predictions, which can enable selective prediction and improve user confidence, is challenging d…