12 papers
A Survey of Toxicity Detection and Mitigation Strategies for Multilingual Language Models
Soham Dan, Himanshu Beniwal, Thomas Hartvigsen
Large language models (LLMs) are increasingly deployed across languages, but their safety behavior remains uneven across linguistic and cultural contexts. This survey synthesizes w…
Sycophancy as a Multilingual Alignment Failure: How Safety Degrades Across Languages, Topics, and Models
Arya Shah, Himanshu Beniwal, Mayank Singh +1
Safety-aligned large language models often exhibit sycophancy, which is the tendency to affirm users' opinions regardless of factual accuracy. Although well-studied in English, its…
DEPART: DEcomposing PARiTy across Multilingual LLMs
Manan Uppadhyay, Prashant Kodali, Pranjal Chitale +3
Multilingual Large Language Models (mLLMs) leaderboards report per-language accuracy but rarely explain why disparities emerge, leaving systemic biases unattributed and offering pr…
Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models
Himanshu Beniwal, Mayank Singh
Large language models frequently generate toxic, hateful, or harmful content, yet existing mitigation methods rely on costly retraining or output-level filtering with no mechanisti…
Beyond Monolingual Assumptions: A Survey of Code-Switched NLP in the Era of Large Language Models across Modalities
Rajvee Sheth, Samridhi Raj Sinha, Mahavir Patil +2
Amidst the rapid advances of large language models (LLMs), most LLMs still struggle with mixed-language inputs, limited Codeswitching (CSW) datasets, and evaluation biases, which h…
One Instruction Does Not Fit All: How Well Do Embeddings Align Personas and Instructions in Low-Resource Indian Languages?
Arya Shah, Himanshu beniwal, Mayank Singh
Aligning multilingual assistants with culturally grounded user preferences is essential for serving India's linguistically diverse population of over one billion speakers across mu…