3 papers
cs.CL2026
TRACES: Proactive Safety Auditing for Multi-Turn LLM Agents via Trajectory-State Modeling
Jiaqian Li, Yanshu Li, Boxuan Zhang +2
LLM agents increasingly operate through multi-turn tool use and environment interaction, where safety risks often emerge from intermediate steps long before they surface in the fin…
cs.CL2026
RTD-Guard: A Black-Box Textual Adversarial Detection Framework via Replacement Token Detection
He Zhu, Yanshu Li, Wen Liu +1
Textual adversarial attacks pose a serious security threat to Natural Language Processing (NLP) systems by introducing imperceptible perturbations that mislead deep learning models…
cs.CR2025
Never compromise with vulnerabilities: a comprehensive survey on AI governance
Yuchu Jiang, Jian Zhao, Yuchen Yuan +64
The rapid advancement of AI has expanded its capabilities across domains, yet introduced critical technical vulnerabilities, such as algorithmic bias and adversarial sensitivity, t…