4 papers
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis
Zhisheng Zheng, Xiaohang Sun, Zhu Liu +5
Recent Text-To-Speech (TTS) systems have achieved strong naturalness and zero-shot voice cloning performance, but fine-grained control of expressive speech at the word or phoneme l…
RosettaSpeech: Zero-Shot Speech-to-Speech Translation without Parallel Speech
Zhisheng Zheng, Xiaohang Sun, Tuan Dinh +8
End-to-end speech-to-speech translation (S2ST) systems typically struggle with a critical data bottleneck: the scarcity of parallel speech-to-speech corpora. To overcome this, we i…
Group-Aware Reinforcement Learning for Output Diversity in Large Language Models
Oron Anschel, Alon Shoshan, Adam Botach +7
Large Language Models (LLMs) often suffer from mode collapse, repeatedly generating the same few completions even when many valid answers exist, limiting their diversity across a w…
LV-MAE: Learning Long Video Representations through Masked-Embedding Autoencoders
Ilan Naiman, Emanuel Ben-Baruch, Oron Anschel +4
In this work, we introduce long-video masked-embedding autoencoders (LV-MAE), a self-supervised learning framework for long video representation. Our approach treats short- and lon…