12 papers
Low-Latency Real-Time Audio Game Commentary System via LLM-Based Parallel Text Generation
Ryota Kawamatsu, Anum Afzal, Yuki Saito +5
We present a low-latency real-time audio game commentary system that generates spoken commentary directly from live gameplay video. In this end-to-end setting, a key bottleneck is…
Top-down string-to-dependency Neural Machine Translation
Shuhei Kondo, Katsuhito Sudoh, Yuji Matsumoto
Most of modern neural machine translation (NMT) models are based on an encoder-decoder framework with an attention mechanism. While they perform well on standard datasets, they can…
Real-Time Generation of Game Video Commentary with Multimodal LLMs: Pause-Aware Decoding Approaches
Anum Afzal, Yuki Saito, Hiroya Takamura +5
Real-time video commentary generation provides textual descriptions of ongoing events in videos. It supports accessibility and engagement in domains such as sports, esports, and li…
An Automatic Quality Metric for Evaluating Simultaneous Interpretation
Mana Makinae, Katsuhito Sudoh, Masaru Yamada +1
Simultaneous interpretation (SI), the translation of one language to another in real time, starts translation before the original speech has finished. Its evaluation needs to consi…
Findings of the IWSLT 2024 Evaluation Campaign
Ibrahim Said Ahmad, Antonios Anastasopoulos, OndÅej Bojar +42
This paper reports on the shared tasks organized by the 21st IWSLT Conference. The shared tasks address 7 scientific challenges in spoken language translation: simultaneous and off…
FuxiTranyu: A Multilingual Large Language Model Trained with Balanced Data
Haoran Sun, Renren Jin, Shaoyang Xu +10
Large language models (LLMs) have demonstrated prowess in a wide range of tasks. However, many LLMs exhibit significant performance discrepancies between high- and low-resource lan…