2 papers
cs.RO2025
EmojiVoice: Towards long-term controllable expressivity in robot speech
Paige TuttösÃ, Shivam Mehta, Zachary Syvenky +3
Humans vary their expressivity when speaking for extended periods to maintain engagement with their listener. Although social robots tend to be deployed with ``expressive'' joyful…
eess.AS2025
Make Some Noise: Towards LLM audio reasoning and generation using sound tokens
Shivam Mehta, Nebojsa Jojic, Hannes Gamper
Integrating audio comprehension and generation into large language models (LLMs) remains challenging due to the continuous nature of audio and the resulting high sampling rates. He…