67 citations · 80 across the 8 of their papers we have counts for
9 papers
Cross-Lingual Ability of Multilingual Masked Language Models: A Study of Language Structure
Yuan Chai, Yaobo Liang, Nan Duan
Multilingual pre-trained language models, such as mBERT and XLM-R, have shown impressive cross-lingual ability. Surprisingly, both of them use multilingual masked language model (M…
Multi-View Document Representation Learning for Open-Domain Dense Retrieval
Shunyu Zhang, Yaobo Liang, Ming Gong +2
Dense retrieval has achieved impressive advances in first-stage retrieval from a large-scale document collection, which is built on bi-encoder architecture to produce single vector…
Discovering Representation Sprachbund For Multilingual Pre-Training
Yimin Fan, Yaobo Liang, Alexandre Muzio +4
Multilingual pre-trained models have demonstrated their effectiveness in many multilingual NLP tasks and enabled zero-shot or few-shot transfer from high-resource languages to low…
Simpson's Bias in NLP Training
Fei Yuan, Longtu Zhang, Huang Bojun +1
In most machine learning tasks, we evaluate a model on a given data population by measuring a population-level metric . Examples of such evaluation metric inclu…
Tag and Correct: Question aware Open Information Extraction with Two-stage Decoding
Martin Kuo, Yaobo Liang, Lei Ji +4
Question Aware Open Information Extraction (Question aware Open IE) takes question and passage as inputs, outputting an answer tuple which contains a subject, a predicate, and one…
XGLUE: A New Benchmark Dataset for Cross-lingual Pre-training, Understanding and Generation
Yaobo Liang, Nan Duan, Yeyun Gong +21
In this paper, we introduce XGLUE, a new benchmark dataset that can be used to train large-scale cross-lingual pre-trained models using multilingual and bilingual corpora and evalu…