3 papers
cs.CR2026
Trace: Unmasking AI Attack Agents Through Terminal Behavior Fingerprinting
Murali Ediga, Sudipta Chattopadhyay
AI-driven penetration testing agents are now capable of autonomously executing attacks within compromised networks. Identifying the model family that controls the active sessions o…
cs.CR2025
Localizing Malicious Outputs from CodeLLM
Mayukh Borana, Junyi Liang, Sai Sathiesh Rajan +1
We introduce FreqRank, a mutation-based defense to localize malicious components in LLM outputs and their corresponding backdoor triggers. FreqRank assumes that the malicious sub-s…
cs.CL2025
HInter: Exposing Hidden Intersectional Bias in Large Language Models
Badr Souani, Ezekiel Soremekun, Mike Papadakis +3
Large Language Models (LLMs) may portray discrimination towards certain individuals, especially those characterized by multiple attributes (aka intersectional bias). Discovering in…