activity
20122020
most citedFinding Linear Structure in Large Datasets with Scalable Canonical Correlation Analysis

35 citations · 65 across the 6 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2020

Unsupervised Bitext Mining and Translation via Self-trained Contextual Embeddings

Phillip Keung, Julian Salazar, Yichao Lu +1

We describe an unsupervised method to create pseudo-parallel corpora for machine translation (MT) from unaligned text. We use multilingual BERT to create source and target sentence…

cs.CL2020

The Multilingual Amazon Reviews Corpus

Phillip Keung, Yichao Lu, György Szarvas +1

We present the Multilingual Amazon Reviews Corpus (MARC), a large-scale collection of Amazon reviews for multilingual text classification. The corpus contains reviews in English, J…

cs.CL2020

Don't Use English Dev: On the Zero-Shot Cross-Lingual Evaluation of Contextual Embeddings

Phillip Keung, Yichao Lu, Julian Salazar +1

Multilingual contextual embeddings have demonstrated state-of-the-art performance in zero-shot cross-lingual transfer learning, where multilingual BERT is fine-tuned on one source…

cs.CL2019

Adversarial Learning with Contextual Embeddings for Zero-resource Cross-lingual Classification and NER

Phillip Keung, Yichao Lu, Vikas Bhardwaj

Contextual word embeddings (e.g. GPT, BERT, ELMo, etc.) have demonstrated state-of-the-art performance on various NLP tasks. Recent work with the multilingual version of BERT has s…

cs.CL2018

A neural interlingua for multilingual machine translation

Yichao Lu, Phillip Keung, Faisal Ladhak +3

We incorporate an explicit neural interlingua into a multilingual encoder-decoder neural machine translation (NMT) architecture. We demonstrate that our model learns a language-ind…

cs.CL201713 cited

A practical approach to dialogue response generation in closed domains

Yichao Lu, Phillip Keung, Shaonan Zhang +2

We describe a prototype dialogue response generation model for the customer service domain at Amazon. The model, which is trained in a weakly supervised fashion, measures the simil…