activity
20182021
most citedA Call for Prudent Choice of Subword Merge Operations in Neural Machine Translation

31 citations · 46 across the 4 of their papers we have counts for

collaborators

9 papers

cs.CL202112 cited

Investigating Failures of Automatic Translation in the Case of Unambiguous Gender

Adithya Renduchintala, Adina Williams

Transformer based models are the modern work horses for neural machine translation (NMT), reaching state of the art across several benchmarks. Despite their impressive accuracy, we…

cs.CL2021

XLEnt: Mining a Large Cross-lingual Entity Dataset with Lexical-Semantic-Phonetic Word Alignment

Ahmed El-Kishky, Adithya Renduchintala, James Cross +2

Cross-lingual named-entity lexica are an important resource to multilingual NLP tasks such as machine translation and cross-lingual wikification. While knowledge bases contain a la…

cs.LG20211 cited

Towards Understanding the Behaviors of Optimal Deep Active Learning Algorithms

Yilun Zhou, Adithya Renduchintala, Xian Li +3

Active learning (AL) algorithms may achieve better performance with fewer data because the model guides the data selection process. While many algorithms have been proposed, there…

cs.CL20212 cited

Quality Estimation without Human-labeled Data

Yi-Lin Tuan, Ahmed El-Kishky, Adithya Renduchintala +3

Quality estimation aims to measure the quality of translated content without access to a reference translation. This is crucial for machine translation systems in real-world scenar…

cs.CL201931 cited

A Call for Prudent Choice of Subword Merge Operations in Neural Machine Translation

Shuoyang Ding, Adithya Renduchintala, Kevin Duh

Most neural machine translation systems are built upon subword units extracted by methods such as Byte-Pair Encoding (BPE) or wordpiece. However, the choice of number of merge oper…

eess.AS2018

Pretraining by Backtranslation for End-to-end ASR in Low-Resource Settings

Matthew Wiesner, Adithya Renduchintala, Shinji Watanabe +3

We explore training attention-based encoder-decoder ASR in low-resource settings. These models perform poorly when trained on small amounts of transcribed speech, in part because t…