2 citations · 2 across the 4 of their papers we have counts for
6 papers
BhashaKritika: Building Synthetic Pretraining Data at Scale for Indic Languages
Guduru Manoj, Neel Prabhanjan Rachamalla, Ashish Kulkarni +8
In the context of pretraining of Large Language Models (LLMs), synthetic data has emerged as an alternative for generating high-quality pretraining data at scale. This is particula…
IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs
Ali Faraz, Akash, Shaharukh Khan +6
Vision-language models (VLMs) have demonstrated impressive generalization across multimodal tasks, yet most evaluation benchmarks remain Western-centric, leaving open questions abo…
Pragyaan: Designing and Curating High-Quality Cultural Post-Training Datasets for Indian Languages
Neel Prabhanjan Rachamalla, Aravind Konakalla, Gautam Rajeev +3
The effectiveness of Large Language Models (LLMs) depends heavily on the availability of high-quality post-training data, particularly instruction-tuning and preference-based examp…
Chitranuvad: Adapting Multi-Lingual LLMs for Multimodal Translation
Shaharukh Khan, Ayush Tarun, Ali Faraz +7
In this work, we provide the system description of our submission as part of the English to Lowres Multimodal Translation Task at the Workshop on Asian Translation (WAT2024). We in…
Krutrim LLM: Multilingual Foundational Model for over a Billion People
Aditya Kallappa, Palash Kamble, Abhinav Ravi +10
India is a diverse society with unique challenges in developing AI systems, including linguistic diversity, oral traditions, data accessibility, and scalability. Existing foundatio…
Chitrarth: Bridging Vision and Language for a Billion People
Shaharukh Khan, Ayush Tarun, Abhinav Ravi +7
Recent multimodal foundation models are primarily trained on English or high resource European language data, which hinders their applicability to other medium and low-resource lan…