activity
20182022
most citedAI4Bharat-IndicNLP Corpus: Monolingual Corpora and Word Embeddings for Indic Languages

43 citations · 71 across the 4 of their papers we have counts for

collaborators

11 papers

cs.CL20225 cited

IndicXNLI: Evaluating Multilingual Inference for Indian Languages

Divyanshu Aggarwal, Vivek Gupta, Anoop Kunchukuttan

While Indic NLP has made rapid advances recently in terms of the availability of corpora and pre-trained models, benchmark datasets on standard NLU tasks are limited. To this end,…

cs.CL2021

An Empirical Investigation of Multi-bridge Multilingual NMT models

Anoop Kunchukuttan

In this paper, we present an extensive investigation of multi-bridge, many-to-many multilingual NMT models (MB-M2M) ie., models trained on non-English language pairs in addition to…

cs.CL2021

Itihasa: A large-scale corpus for Sanskrit to English translation

Rahul Aralikatte, Miryam de Lhoneux, Anoop Kunchukuttan +1

This work introduces Itihasa, a large-scale translation dataset containing 93,000 pairs of Sanskrit shlokas and their English translations. The shlokas are extracted from two India…

cs.CL202043 cited

AI4Bharat-IndicNLP Corpus: Monolingual Corpora and Word Embeddings for Indic Languages

Anoop Kunchukuttan, Divyanshu Kakwani, Satish Golla +4

We present the IndicNLP corpus, a large-scale, general-domain corpus containing 2.7 billion words for 10 Indian languages from two language families. We share pre-trained word embe…

cs.CL2020

Learning Geometric Word Meta-Embeddings

Pratik Jawanpuria, N T V Satya Dev, Anoop Kunchukuttan +1

We propose a geometric framework for learning meta-embeddings of words from different embedding sources. Our framework transforms the embeddings into a common latent space, where,…

cs.CL2020

Utilizing Language Relatedness to improve Machine Translation: A Case Study on Languages of the Indian Subcontinent

Anoop Kunchukuttan, Pushpak Bhattacharyya

In this work, we present an extensive study of statistical machine translation involving languages of the Indian subcontinent. These languages are related by genetic and contact re…