68 citations · 219 across the 24 of their papers we have counts for
29 papers · 1 filter
Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis
Sanyuan Chen, Min-Jae Hwang, Sho Inoue +12
We present Alignment-Free Text-Audiobox (Text-AB), a unified framework for high-quality voice dubbing and full-duplex dialogue synthesis. Building on a Diffusion Transformer traine…
SeamlessM4T: Massively Multilingual & Multimodal Machine Translation
Seamless Communication, Loïc Barrault, Yu-An Chung +65
What does it take to create the Babel Fish, a tool that can help individuals translate speech between any two languages? While recent breakthroughs in text-based models have pushed…
Multilingual Speech-to-Speech Translation into Multiple Target Languages
Hongyu Gong, Ning Dong, Sravya Popuri +3
Speech-to-speech translation (S2ST) enables spoken communication between people talking in different languages. Despite a few studies on multilingual S2ST, their focus is the multi…
Hybrid Transducer and Attention based Encoder-Decoder Modeling for Speech-to-Text Tasks
Yun Tang, Anna Y. Sun, Hirofumi Inaguma +5
Transducer and Attention based Encoder-Decoder (AED) are two widely used frameworks for speech-to-text tasks. They are designed for different purposes and each has its own benefits…
MuAViC: A Multilingual Audio-Visual Corpus for Robust Speech Recognition and Robust Speech-to-Text Translation
Mohamed Anwar, Bowen Shi, Vedanuj Goswami +3
We introduce MuAViC, a multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation providing 1200 hours of audio-visual speech in 9 languag…
Speech-to-Speech Translation For A Real-world Unwritten Language
Peng-Jen Chen, Kevin Tran, Yilin Yang +13
We study speech-to-speech translation (S2ST) that translates speech from one language into another language and focuses on building systems to support languages without standard te…