12 citations · 30 across the 8 of their papers we have counts for
11 papers
JParaCrawl v3.0: A Large-scale English-Japanese Parallel Corpus
Makoto Morishita, Katsuki Chousa, Jun Suzuki +1
Most current machine translation models are mainly trained with parallel corpora, and their translation accuracy largely depends on the quality and quantity of the corpora. Althoug…
Bilingual Text Extraction as Reading Comprehension
Katsuki Chousa, Masaaki Nagata, Masaaki Nishino
In this paper, we propose a method to extract bilingual texts automatically from noisy parallel corpora by framing the problem as a token-level span prediction, such as SQuAD-style…
A Supervised Word Alignment Method based on Cross-Language Span Prediction using Multilingual BERT
Masaaki Nagata, Chousa Katsuki, Masaaki Nishino
We present a novel supervised word alignment method based on cross-language span prediction. We first formalize a word alignment problem as a collection of independent predictions…
JParaCrawl: A Large Scale Web-Based English-Japanese Parallel Corpus
Makoto Morishita, Jun Suzuki, Masaaki Nagata
Recent machine translation algorithms mainly rely on parallel corpora. However, since the availability of parallel corpora remains limited, only some resource-rich language pairs c…
NTT's Machine Translation Systems for WMT19 Robustness Task
Soichiro Murakami, Makoto Morishita, Tsutomu Hirao +1
This paper describes NTT's submission to the WMT19 robustness task. This task mainly focuses on translating noisy text (e.g., posts on Twitter), which presents different difficulti…
Character n-gram Embeddings to Improve RNN Language Models
Sho Takase, Jun Suzuki, Masaaki Nagata
This paper proposes a novel Recurrent Neural Network (RNN) language model that takes advantage of character information. We focus on character n-grams based on research in the fiel…