2 papers
cs.CL2024
Enhancing Multilingual Voice Toxicity Detection with Speech-Text Alignment
Joseph Liu, Mahesh Kumar Nandwana, Janne Pylkkönen +2
Toxicity classification for voice heavily relies on the semantic content of speech. We propose a novel framework that utilizes cross-modal learning to integrate the semantic embedd…
cs.LG2024
Diffusion Synthesizer for Efficient Multilingual Speech to Speech Translation
Nameer Hirschkind, Xiao Yu, Mahesh Kumar Nandwana +9
We introduce DiffuseST, a low-latency, direct speech-to-speech translation system capable of preserving the input speaker's voice zero-shot while translating from multiple source l…