4 papers
Minionese: Comprehensive Benchmark and Mechanistic Study of Multilingual LLM Safety
Chigozirim Ifebi, Brent Kong, Ayushi Mehrotra
Safety alignment in large language models remains brittle across languages: prompts reliably refused in English can elicit harmful compliance in non-English and low-resource settin…
H-Sets: Hessian-Guided Discovery of Set-Level Feature Interactions in Image Classifiers
Ayushi Mehrotra, Dipkamal Bhusal, Michael Clifford +1
Feature attribution methods explain the predictions of deep neural networks by assigning importance scores to individual input features. However, most existing methods focus solely…
Towards Realistic Guarantees: A Probabilistic Certificate for SmoothLLM
Adarsh Kumarappan, Ayushi Mehrotra
The SmoothLLM defense provides a certification guarantee against jailbreaking attacks, but it relies on a strict "k-unstable" assumption that rarely holds in practice. This strong…
Concept-Based Masking: A Patch-Agnostic Defense Against Adversarial Patch Attacks
Ayushi Mehrotra, Derek Peng, Dipkamal Bhusal +1
Adversarial patch attacks pose a practical threat to deep learning models by forcing targeted misclassifications through localized perturbations, often realized in the physical wor…