67 citations · 76 across the 14 of their papers we have counts for
19 papers · 1 filter
R-WoM: Retrieval-augmented World Model For Computer-use Agents
Kai Mei, Jiang Guo, Shuaichen Chang +4
Large Language Models (LLMs) can serve as world models to enhance agent decision-making in digital environments by simulating future states and predicting action outcomes, potentia…
Zero-resource Speech Translation and Recognition with LLMs
Karel Mundnich, Xing Niu, Prashant Mathur +10
Despite recent advancements in speech processing, zero-resource speech translation (ST) and automatic speech recognition (ASR) remain challenging problems. In this work, we propose…
Findings of the IWSLT 2024 Evaluation Campaign
Ibrahim Said Ahmad, Antonios Anastasopoulos, Ondřej Bojar +42
This paper reports on the shared tasks organized by the 21st IWSLT Conference. The shared tasks address 7 scientific challenges in spoken language translation: simultaneous and off…
M3T: A New Benchmark Dataset for Multi-Modal Document-Level Machine Translation
Benjamin Hsu, Xiaoyu Liu, Huayang Li +6
Document translation poses a challenge for Neural Machine Translation (NMT) systems. Most document-level NMT systems rely on meticulously curated sentence-level parallel data, assu…
SpeechVerse: A Large-scale Generalizable Audio Language Model
Nilaksh Das, Saket Dingliwal, Srikanth Ronanki +14
Large language models (LLMs) have shown incredible proficiency in performing tasks that require semantic understanding of natural language instructions. Recently, many works have f…
End-to-End Single-Channel Speaker-Turn Aware Conversational Speech Translation
Juan Zuluaga-Gomez, Zhaocheng Huang, Xing Niu +5
Conventional speech-to-text translation (ST) systems are trained on single-speaker utterances, and they may not generalize to real-life scenarios where the audio contains conversat…