most citedMonotonic Multihead Attention

68 citations · 129 across the 6 of their papers we have counts for

collaborators
Showing cs.CLShow all

11 papers · 1 filter

cs.CL202314 cited

SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

Seamless Communication, Loïc Barrault, Yu-An Chung +65

What does it take to create the Babel Fish, a tool that can help individuals translate speech between any two languages? While recent breakthroughs in text-based models have pushed…

cs.CL20231 cited

Hybrid Transducer and Attention based Encoder-Decoder Modeling for Speech-to-Text Tasks

Yun Tang, Anna Y. Sun, Hirofumi Inaguma +5

Transducer and Attention based Encoder-Decoder (AED) are two widely used frameworks for speech-to-text tasks. They are designed for different purposes and each has its own benefits…

cs.CL202018 cited

SimulMT to SimulST: Adapting Simultaneous Text Translation to End-to-End Simultaneous Speech Translation

Xutai Ma, Juan Pino, Philipp Koehn

Simultaneous text translation and end-to-end speech translation have recently made great progress but little work has combined these tasks together. We investigate how to adapt sim…

cs.CL20202 cited

Streaming Simultaneous Speech Translation with Augmented Memory Transformer

Xutai Ma, Yongqiang Wang, Mohammad Javad Dousti +2

Transformer-based models have achieved state-of-the-art performance on speech translation tasks. However, the model architecture is not efficient enough for streaming scenarios sin…

cs.CL2020

A General Multi-Task Learning Framework to Leverage Text Data for Speech to Text Tasks

Yun Tang, Juan Pino, Changhan Wang +2

Attention-based sequence-to-sequence modeling provides a powerful and elegant solution for applications that need to map one sequence to a different sequence. Its success heavily r…

cs.CL20206 cited

SimulEval: An Evaluation Toolkit for Simultaneous Translation

Xutai Ma, Mohammad Javad Dousti, Changhan Wang +2

Simultaneous translation on both text and speech focuses on a real-time and low-latency scenario where the model starts translating before reading the complete source input. Evalua…