8 papers
Forecasting Side Effects of Activation Steering
Chong Yong Ong, Alson Wei Jie Sim, Peixin Zhang +1
Activation steering modifies a language model by adding a learned direction to its hidden activations, enabling targeted behavioral changes without retraining. While effective, ste…
Efficient and Universal Watermarking for LLM-Generated Code Detection
Boquan Li, Zirui Fu, Mengdi Zhang +3
Large language models (LLMs) have significantly enhanced the usability of AI-generated code, providing effective assistance to programmers. This advancement also raises ethical and…
DDOR: Delta Debugging for Explainable Overrefusal Testing and Repair
Qinyan Zhou, Peixin Zhang, Jun Sun +2
While safety alignment and guardrails help large language models (LLMs) avoid harmful outputs, they can also induce overrefusal, i.e., unwarranted rejection of benign queries that…
ClawGuard: A Runtime Security Framework for Tool-Augmented LLM Agents Against Indirect Prompt Injection
Wei Zhao, Zhe Li, Peixin Zhang +1
Tool-augmented Large Language Model (LLM) agents have demonstrated impressive capabilities in automating complex, multi-step real-world tasks, yet remain vulnerable to indirect pro…
The Salami Slicing Threat: Exploiting Cumulative Risks in LLM Systems
Yihao Zhang, Kai Wang, Jiangrong Wu +7
Large Language Models (LLMs) face prominent security risks from jailbreaking, a practice that manipulates models to bypass built-in security constraints and generate unethical or u…
NuHF Claw: A Risk Constrained Cognitive Agent Framework for Human Centered Procedure Support in Digital Nuclear Control Rooms
Xingyu Xiao, Jiejuan Tong, Jun Sun +4
The rapid digitization of nuclear power plant main control rooms has fundamentally reshaped operator interaction patterns, introducing complex soft-control behaviors and elevated c…