2 papers
cs.CL2026
Hessian-Enhanced Token Attribution (HETA): Interpreting Autoregressive LLMs
Vishal Pramanik, Maisha Maliha, Nathaniel D. Bastian +1
Attribution methods seek to explain language model predictions by quantifying the contribution of input tokens to generated outputs. However, most existing techniques are designed…
cs.CR2026
Jailbreaking the Matrix: Nullspace Steering for Controlled Model Subversion
Vishal Pramanik, Maisha Maliha, Susmit Jha +1
Large language models remain vulnerable to jailbreak attacks -- inputs designed to bypass safety mechanisms and elicit harmful responses -- despite advances in alignment and instru…