6 citations · 10 across the 3 of their papers we have counts for
4 papers
TEASEL: A Transformer-Based Speech-Prefixed Language Model
Mehdi Arjmand, Mohammad Javad Dousti, Hadi Moradi
Multimodal language analysis is a burgeoning field of NLP that aims to simultaneously model a speaker's words, acoustical annotations, and facial expressions. In this area, lexicon…
Streaming Simultaneous Speech Translation with Augmented Memory Transformer
Xutai Ma, Yongqiang Wang, Mohammad Javad Dousti +2
Transformer-based models have achieved state-of-the-art performance on speech translation tasks. However, the model architecture is not efficient enough for streaming scenarios sin…
SimulEval: An Evaluation Toolkit for Simultaneous Translation
Xutai Ma, Mohammad Javad Dousti, Changhan Wang +2
Simultaneous translation on both text and speech focuses on a real-time and low-latency scenario where the model starts translating before reading the complete source input. Evalua…
Self-Training for End-to-End Speech Translation
Juan Pino, Qiantong Xu, Xutai Ma +2
One of the main challenges for end-to-end speech translation is data scarcity. We leverage pseudo-labels generated from unlabeled audio by a cascade and an end-to-end speech transl…