6 papers
Redirected, Not Removed: Task-Dependent Stereotyping Reveals the Limits of LLM Alignments
Divyanshu Kumar, Ishita Gupta, Nitin Aravind Birur +3
How biased is a language model? The answer depends on how you ask. A model that refuses to choose between castes for a leadership role will, in a fill-in-the-blank task, reliably a…
SocioEval: A Template-Based Framework for Evaluating Socioeconomic Status Bias in Foundation Models
Divyanshu Kumar, Ishita Gupta, Nitin Aravind Birur +3
As Large Language Models (LLMs) increasingly power decision-making systems across critical domains, understanding and mitigating their biases becomes essential for responsible AI d…
Quantifying CBRN Risk in Frontier Models
Divyanshu Kumar, Nitin Aravind Birur, Tanay Baswa +2
Frontier Large Language Models (LLMs) pose unprecedented dual-use risks through the potential proliferation of chemical, biological, radiological, and nuclear (CBRN) weapons knowle…
Beyond Text: Multimodal Jailbreaking of Vision-Language and Audio Models through Perceptually Simple Transformations
Divyanshu Kumar, Shreyas Jena, Nitin Aravind Birur +3
Multimodal large language models (MLLMs) have achieved remarkable progress, yet remain critically vulnerable to adversarial attacks that exploit weaknesses in cross-modal processin…
Beyond Western Politics: Cross-Cultural Benchmarks for Evaluating Partisan Associations in LLMs
Divyanshu Kumar, Ishita Gupta, Nitin Aravind Birur +3
Partisan bias in LLMs has been evaluated to assess political leanings, typically through a broad lens and largely in Western contexts. We move beyond identifying general leanings t…
No Free Lunch with Guardrails
Divyanshu Kumar, Nitin Aravind Birur, Tanay Baswa +2
As large language models (LLMs) and generative AI become widely adopted, guardrails have emerged as a key tool to ensure their safe use. However, adding guardrails isn't without tr…