2 papers
cs.LG2026
Towards Understanding the Robustness of Sparse Autoencoders
Ahson Saiyed, Sabrina Sadiekh, Chirag Agarwal
Large Language Models (LLMs) remain vulnerable to optimization-based jailbreak attacks that exploit internal gradient structure. While Sparse Autoencoders (SAEs) are widely used fo…
cs.LG2025
Analyzing Memorization in Large Language Models through the Lens of Model Attribution
Tarun Ram Menta, Susmit Agrawal, Chirag Agarwal
Large Language Models (LLMs) are prevalent in modern applications but often memorize training data, leading to privacy breaches and copyright issues. Existing research has mainly f…