4 citations · 5 across the 9 of their papers we have counts for
4 papers · 1 filter
Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models
Bhuvan Koduru, Dareen Safar B Alharthi, Rita Singh +1
Audio language models are designed to understand speech, yet it remains unclear whether they capture how something is said beyond what is said. We present a mechanistic analysis of…
RIVET: Robust Idempotent Voice Attribute Editing
Dareen Alharthi, Bhuvan Koduru, Rita Singh +1
Voice attribute editing models modify characteristics such as age and gender while preserving speaker identity. In large-scale speech datasets, however, attribute annotations are o…
Domain Adaptation for Contrastive Audio-Language Models
Soham Deshmukh, Rita Singh, Bhiksha Raj
Audio-Language Models (ALM) aim to be general-purpose audio models by providing zero-shot capabilities at test time. The zero-shot performance of ALM improves by using suitable tex…
Prompting Audios Using Acoustic Properties For Emotion Representation
Hira Dhamyal, Benjamin Elizalde, Soham Deshmukh +3
Emotions lie on a continuum, but current models treat emotions as a finite valued discrete variable. This representation does not capture the diversity in the expression of emotion…