Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
The Safety Relay in Roleplay Jailbreaks: A Component-Resolved Causal Analysis of Harm Recognition and Refusal
Md Mokarram Chowdhury, Ernie Chang, Yang Li
Large language models are trained to follow instructions while refusing harmful requests. Jailbreaks exploit this balance to elicit content a model would ordinarily reject. Rolepla…
cs.LG2025
IMPACT: Importance-Aware Activation Space Reconstruction
Md Mokarram Chowdhury, Daniel Agyei Asante, Ernie Chang +1
Large language models (LLMs) achieve strong performance across diverse domains but remain difficult to deploy in resource-constrained environments due to their size. Low-rank compr…