6 papers
Mathematics Isn't Culture-Free: Probing Cultural Gaps via Entity and Scenario Perturbations
Aditya Tomar, Nihar Ranjan Sahoo, Ashish Mittal +2
Although mathematics is often considered culturally neutral, the way mathematical problems are presented can carry implicit cultural context. Existing benchmarks like GSM8K are pre…
BharatBBQ: A Multilingual Bias Benchmark for Question Answering in the Indian Context
Aditya Tomar, Nihar Ranjan Sahoo, Pushpak Bhattacharyya
Evaluating social biases in language models (LMs) is crucial for ensuring fairness and minimizing the reinforcement of harmful stereotypes in AI systems. Existing benchmarks, such…
Pretraining Language Models Using Translationese
Meet Doshi, Raj Dabre, Pushpak Bhattacharyya
In this paper, we explore the utility of translationese as synthetic data created using machine translation for pre-training language models (LMs) for low-resource languages (LRLs)…
Stereotype Detection as a Catalyst for Enhanced Bias Detection: A Multi-Task Learning Approach
Aditya Tomar, Rudra Murthy, Pushpak Bhattacharyya
Bias and stereotypes in language models can cause harm, especially in sensitive areas like content moderation and decision-making. This paper addresses bias and stereotype detectio…
Reconsidering SMT Over NMT for Closely Related Languages: A Case Study of Persian-Hindi Pair
Waisullah Yousofi, Pushpak Bhattacharyya
This paper demonstrates that Phrase-Based Statistical Machine Translation (PBSMT) can outperform Transformer-based Neural Machine Translation (NMT) in moderate-resource scenarios,…
Are Language Models Agnostic to Linguistically Grounded Perturbations? A Case Study of Indic Languages
Poulami Ghosh, Raj Dabre, Pushpak Bhattacharyya
Pre-trained language models (PLMs) are known to be susceptible to perturbations to the input text, but existing works do not explicitly focus on linguistically grounded attacks, wh…