3 papers
cs.CR2026
MLingualFC: Evaluating Jailbreak Vulnerabilities in Multilingual Vision-Language Models
Rishabh Makwana, Mamta, Deeksha Varshney +1
Vision-Language Models (VLMs) have demonstrated strong performance across multimodal tasks, yet their safety robustness remains an open challenge. While prior work has shown that s…
cs.CL2026
Indic-TunedLens: Interpreting Multilingual Models in Indian Languages
Mihir Panchal, Deeksha Varshney, Mamta +1
Multilingual large language models (LLMs) are increasingly deployed in linguistically diverse regions like India, yet most interpretability tools remain tailored to English. Prior…
cs.CL2025
Concept-Based Interpretability for Toxicity Detection
Samarth Garg, Divya Singh, Deeksha Varshney +1
The rise of social networks has not only facilitated communication but also allowed the spread of harmful content. Although significant advances have been made in detecting toxic l…