2 papers
cs.CL2026
From Concept-Aligned Tokens to Vulnerable Features: Mechanistic Localization of Jailbreaks
Nilanjana Das, Mathew Dawit, Aman Chadha +1
Jailbreak attacks expose a persistent failure mode in safety-aligned LLMs: models can be pushed into harmful behavior, but the internal representations enabling this shift remain p…
cs.CR2023
Change Management using Generative Modeling on Digital Twins
Nilanjana Das, Anantaa Kotal, Daniel Roseberry +1
A key challenge faced by small and medium-sized business entities is securely managing software updates and changes. Specifically, with rapidly evolving cybersecurity threats, chan…