activity
20172022
most citedSource-side Prediction for Neural Headline Generation

12 citations · 30 across the 8 of their papers we have counts for

collaborators

11 papers

cs.CL2022

JParaCrawl v3.0: A Large-scale English-Japanese Parallel Corpus

Makoto Morishita, Katsuki Chousa, Jun Suzuki +1

Most current machine translation models are mainly trained with parallel corpora, and their translation accuracy largely depends on the quality and quantity of the corpora. Althoug…

cs.CL2020

Bilingual Text Extraction as Reading Comprehension

Katsuki Chousa, Masaaki Nagata, Masaaki Nishino

In this paper, we propose a method to extract bilingual texts automatically from noisy parallel corpora by framing the problem as a token-level span prediction, such as SQuAD-style…

cs.CL20206 cited

A Supervised Word Alignment Method based on Cross-Language Span Prediction using Multilingual BERT

Masaaki Nagata, Chousa Katsuki, Masaaki Nishino

We present a novel supervised word alignment method based on cross-language span prediction. We first formalize a word alignment problem as a collection of independent predictions…

cs.CL2019

JParaCrawl: A Large Scale Web-Based English-Japanese Parallel Corpus

Makoto Morishita, Jun Suzuki, Masaaki Nagata

Recent machine translation algorithms mainly rely on parallel corpora. However, since the availability of parallel corpora remains limited, only some resource-rich language pairs c…

cs.CL2019

NTT's Machine Translation Systems for WMT19 Robustness Task

Soichiro Murakami, Makoto Morishita, Tsutomu Hirao +1

This paper describes NTT's submission to the WMT19 robustness task. This task mainly focuses on translating noisy text (e.g., posts on Twitter), which presents different difficulti…

cs.CL20192 cited

Character n-gram Embeddings to Improve RNN Language Models

Sho Takase, Jun Suzuki, Masaaki Nagata

This paper proposes a novel Recurrent Neural Network (RNN) language model that takes advantage of character information. We focus on character n-grams based on research in the fiel…