43 citations · 71 across the 4 of their papers we have counts for
11 papers
IndicXNLI: Evaluating Multilingual Inference for Indian Languages
Divyanshu Aggarwal, Vivek Gupta, Anoop Kunchukuttan
While Indic NLP has made rapid advances recently in terms of the availability of corpora and pre-trained models, benchmark datasets on standard NLU tasks are limited. To this end,…
An Empirical Investigation of Multi-bridge Multilingual NMT models
Anoop Kunchukuttan
In this paper, we present an extensive investigation of multi-bridge, many-to-many multilingual NMT models (MB-M2M) ie., models trained on non-English language pairs in addition to…
Itihasa: A large-scale corpus for Sanskrit to English translation
Rahul Aralikatte, Miryam de Lhoneux, Anoop Kunchukuttan +1
This work introduces Itihasa, a large-scale translation dataset containing 93,000 pairs of Sanskrit shlokas and their English translations. The shlokas are extracted from two India…
AI4Bharat-IndicNLP Corpus: Monolingual Corpora and Word Embeddings for Indic Languages
Anoop Kunchukuttan, Divyanshu Kakwani, Satish Golla +4
We present the IndicNLP corpus, a large-scale, general-domain corpus containing 2.7 billion words for 10 Indian languages from two language families. We share pre-trained word embe…
Learning Geometric Word Meta-Embeddings
Pratik Jawanpuria, N T V Satya Dev, Anoop Kunchukuttan +1
We propose a geometric framework for learning meta-embeddings of words from different embedding sources. Our framework transforms the embeddings into a common latent space, where,…
Utilizing Language Relatedness to improve Machine Translation: A Case Study on Languages of the Indian Subcontinent
Anoop Kunchukuttan, Pushpak Bhattacharyya
In this work, we present an extensive study of statistical machine translation involving languages of the Indian subcontinent. These languages are related by genetic and contact re…