5 citations · 16 across the 6 of their papers we have counts for
6 papers
Synthetic Data Generation and Joint Learning for Robust Code-Mixed Translation
Kartik Kartik, Sanjana Soni, Anoop Kunchukuttan +2
The widespread online communication in a modern multilingual world has provided opportunities to blend more than one language (aka code-mixed language) in a single utterance. This…
Airavata: Introducing Hindi Instruction-tuned LLM
Jay Gala, Thanmay Jayakumar, Jaavid Aktar Husain +8
We announce the initial release of "Airavata," an instruction-tuned LLM for Hindi. Airavata was created by fine-tuning OpenHathi with diverse, instruction-tuning Hindi datasets to…
Evaluating Inter-Bilingual Semantic Parsing for Indian Languages
Divyanshu Aggarwal, Vivek Gupta, Anoop Kunchukuttan
Despite significant progress in Natural Language Generation for Indian languages (IndicNLP), there is a lack of datasets around complex structured tasks such as semantic parsing. O…
Effectiveness of Mining Audio and Text Pairs from Public Data for Improving ASR Systems for Low-Resource Languages
Kaushal Santosh Bhogale, Abhigyan Raman, Tahir Javed +4
End-to-end (E2E) models have become the default choice for state-of-the-art speech recognition systems. Such models are trained on large amounts of labelled data, which are often n…
Faster decoding for subword level Phrase-based SMT between related languages
Anoop Kunchukuttan, Pushpak Bhattacharyya
A common and effective way to train translation systems between related languages is to consider sub-word level basic units. However, this increases the length of the sentences res…
Orthographic Syllable as basic unit for SMT between Related Languages
Anoop Kunchukuttan, Pushpak Bhattacharyya
We explore the use of the orthographic syllable, a variable-length consonant-vowel sequence, as a basic unit of translation between related languages which use abugida or alphabeti…