1 citations · 1 across the 1 of their papers we have counts for
3 papers
Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective
Bhavik Chandna, Zubair Bashir, Procheta Sen
Large Language Models (LLMs) are known to exhibit social, demographic, and gender biases, often as a consequence of the data on which they are trained. In this work, we adopt a mec…
XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content
Vadivel Abishethvarman, Bhavik Chandna, Pratik Jalan +1
Large Language Models (LLMs) can generate content spanning ideological rhetoric to explicit instructions for violence. However, existing safety evaluations often rely on simplistic…
ExtremeAIGC: Benchmarking LMM Vulnerability to AI-Generated Extremist Content
Bhavik Chandna, Mariam Aboujenane, Usman Naseem
Large Multimodal Models (LMMs) are increasingly vulnerable to AI-generated extremist content, including photorealistic images and text, which can be used to bypass safety mechanism…