activity
20172022
most citedText Compression-aided Transformer Encoding

51 citations · 87 across the 14 of their papers we have counts for

collaborators
Showing cs.CLShow all

25 papers · 1 filter

cs.CL2022

Language Model Pre-training on True Negatives

Zhuosheng Zhang, Hai Zhao, Masao Utiyama +1

Discriminative pre-trained language models (PLMs) learn to predict original texts from intentionally corrupted ones. Taking the former text as positive and the latter as negative s…

cs.CL2022

Extending the Subwording Model of Multilingual Pretrained Models for New Languages

Kenji Imamura, Eiichiro Sumita

Multilingual pretrained models are effective for machine translation and cross-lingual processing because they contain multiple languages in one model. However, they are pretrained…

cs.CL2021

Smoothing Dialogue States for Open Conversational Machine Reading

Zhuosheng Zhang, Siru Ouyang, Hai Zhao +2

Conversational machine reading (CMR) requires machines to communicate with humans through multi-turn interactions between two salient dialogue states of decision making and questio…

cs.CL20215 cited

YANMTT: Yet Another Neural Machine Translation Toolkit

Raj Dabre, Eiichiro Sumita

In this paper we present our open-source neural machine translation (NMT) toolkit called "Yet Another Neural Machine Translation Toolkit" abbreviated as YANMTT which is built on to…

cs.CL20213 cited

Cross-lingual Transferring of Pre-trained Contextualized Language Models

Zuchao Li, Kevin Parnow, Hai Zhao +4

Though the pre-trained contextualized language model (PrLM) has made a significant impact on NLP, training PrLMs in languages other than English can be impractical for two reasons:…

cs.CL2021

User-Generated Text Corpus for Evaluating Japanese Morphological Analysis and Lexical Normalization

Shohei Higashiyama, Masao Utiyama, Taro Watanabe +1

Morphological analysis (MA) and lexical normalization (LN) are both important tasks for Japanese user-generated text (UGT). To evaluate and compare different MA/LN systems, we have…