2 papers
cs.CR2026
MLingualFC: Evaluating Jailbreak Vulnerabilities in Multilingual Vision-Language Models
Rishabh Makwana, Mamta, Deeksha Varshney +1
Vision-Language Models (VLMs) have demonstrated strong performance across multimodal tasks, yet their safety robustness remains an open challenge. While prior work has shown that s…
cs.CL2025
Concept-Based Interpretability for Toxicity Detection
Samarth Garg, Divya Singh, Deeksha Varshney +1
The rise of social networks has not only facilitated communication but also allowed the spread of harmful content. Although significant advances have been made in detecting toxic l…