collaborators

10 papers

cs.LG2026

Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks

Nobin Sarwar, Shubhashis Roy Dipta, Zheyuan Liu +1

With the growing adoption of VLMs, DMs, LLMs, and AFMs, these multimodal foundation models can inadvertently encode sensitive, copyrighted, biased, or unsafe cross-modal associatio…

cs.SD2026

AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling

Jiacheng Shi, Hongfei Du, Xinyuan Song +3

Neural speech codecs provide discrete representations for speech language models, but emotional cues are often degraded during quantization. Existing codecs mainly optimize acousti…

eess.AS2026

Conversational Speech Naturalness Predictor

Anfeng Xu, Yashesh Gaur, Naoyuki Kanda +6

Evaluation of conversational naturalness is essential for developing human-like speech agents. However, existing speech naturalness predictors are often designed to assess utteranc…

cs.SD2026

GSRM: Generative Speech Reward Model for Speech RLHF

Maohao Shen, Tejas Jayashankar, Osama Hanna +10

Recent advances in speech language models, such as GPT-4o Voice Mode and Gemini Live, have demonstrated promising speech generation capabilities. Nevertheless, the aesthetic natura…

eess.AS2026

T-Mimi: A Transformer-based Mimi Decoder for Real-Time On-Phone TTS

Haibin Wu, Bach Viet Do, Naveen Suda +10

Neural audio codecs provide promising acoustic features for speech synthesis, with representative streaming codecs like Mimi providing high-quality acoustic features for real-time…

eess.AS2025

Improving Practical Aspects of End-to-End Multi-Talker Speech Recognition for Online and Offline Scenarios

Aswin Shanmugam Subramanian, Amit Das, Naoyuki Kanda +3

We extend the frameworks of Serialized Output Training (SOT) to address practical needs of both streaming and offline automatic speech recognition (ASR) applications. Our approach…