3 papers
cs.CL2026
GradShield: Alignment Preserving Finetuning
Zhanhao Hu, Xiao Huang, Patrick Mendoza +4
Large Language Models (LLMs) pose a significant risk of safety misalignment after finetuning, as models can be compromised by both explicitly and implicitly harmful data. Even some…
cs.CR2026
Preventing Prompt Injection with Type-Directed Privilege Separation
Dennis Jacob, Emad Alghamdi, Zhanhao Hu +2
Modern language models have enabled the development of agentic systems that achieve strong performance on reasoning-intensive tasks. Unfortunately, this has come with a security co…
cs.AI2025
Bridging the Data Provenance Gap Across Text, Speech and Video
Shayne Longpre, Nikhil Singh, Manuel Cherep +40
Progress in AI is driven largely by the scale and quality of training data. Despite this, there is a deficit of empirical analysis examining the attributes of well-established data…