6 citations · 11 across the 4 of their papers we have counts for
6 papers · 1 filter
JParaCrawl v3.0: A Large-scale English-Japanese Parallel Corpus
Makoto Morishita, Katsuki Chousa, Jun Suzuki +1
Most current machine translation models are mainly trained with parallel corpora, and their translation accuracy largely depends on the quality and quantity of the corpora. Althoug…
Input Augmentation Improves Constrained Beam Search for Neural Machine Translation: NTT at WAT 2021
Katsuki Chousa, Makoto Morishita
This paper describes our systems that were submitted to the restricted translation task at WAT 2021. In this task, the systems are required to output translated sentences that cont…
Bilingual Text Extraction as Reading Comprehension
Katsuki Chousa, Masaaki Nagata, Masaaki Nishino
In this paper, we propose a method to extract bilingual texts automatically from noisy parallel corpora by framing the problem as a token-level span prediction, such as SQuAD-style…
A Supervised Word Alignment Method based on Cross-Language Span Prediction using Multilingual BERT
Masaaki Nagata, Chousa Katsuki, Masaaki Nishino
We present a novel supervised word alignment method based on cross-language span prediction. We first formalize a word alignment problem as a collection of independent predictions…
Simultaneous Neural Machine Translation using Connectionist Temporal Classification
Katsuki Chousa, Katsuhito Sudoh, Satoshi Nakamura
Simultaneous machine translation is a variant of machine translation that starts the translation process before the end of an input. This task faces a trade-off between translation…
Training Neural Machine Translation using Word Embedding-based Loss
Katsuki Chousa, Katsuhito Sudoh, Satoshi Nakamura
In neural machine translation (NMT), the computational cost at the output layer increases with the size of the target-side vocabulary. Using a limited-size vocabulary instead may c…