3 papers
cs.SD2026
Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models
Bhuvan Koduru, Dareen Safar B Alharthi, Rita Singh +1
Audio language models are designed to understand speech, yet it remains unclear whether they capture how something is said beyond what is said. We present a mechanistic analysis of…
cs.SD2026
RIVET: Robust Idempotent Voice Attribute Editing
Dareen Alharthi, Bhuvan Koduru, Rita Singh +1
Voice attribute editing models modify characteristics such as age and gender while preserving speaker identity. In large-scale speech datasets, however, attribute annotations are o…
cs.CL2024
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
Roshan Sharma, Suwon Shon, Mark Lindsey +3
Reference summaries for abstractive speech summarization require human annotation, which can be performed by listening to an audio recording or by reading textual transcripts of th…