3 papers
cs.CL2025
ELR-1000: A Community-Generated Dataset for Endangered Indic Indigenous Languages
Neha Joshi, Pamir Gogoi, Aasim Mirza +7
We present a culturally-grounded multimodal dataset of 1,060 traditional recipes crowdsourced from rural communities across remote regions of Eastern India, spanning 10 endangered…
cs.CY2025
What's Not on the Plate? Rethinking Food Computing through Indigenous Indian Datasets
Pamir Gogoi, Neha Joshi, Ayushi Pandey +6
This paper presents a multimodal dataset of 1,000 indigenous recipes from remote regions of India, collected through a participatory model involving first-time digital workers from…
cs.CL2024
PARIKSHA: A Large-Scale Investigation of Human-LLM Evaluator Agreement on Multilingual and Multi-Cultural Data
Ishaan Watts, Varun Gumma, Aditya Yadavalli +3
Evaluation of multilingual Large Language Models (LLMs) is challenging due to a variety of factors -- the lack of benchmarks with sufficient linguistic diversity, contamination of…