7 papers
When Words Are Safe But Actions Kill: Probing Physical Jailbreak Beyond Textual Jailbreak in Hidden-State Risk Space
Weimeng Wang, Ziqiang Wang, Zihang Zhan +3
Large language models (LLMs) increasingly serve as high-level planners for embodied agents, where linguistically benign instructions can become unsafe once grounded in the physical…
SafeSearch: Automated Red-Teaming of LLM-Based Search Agents
Jianshuo Dong, Sheng Guo, Hao Wang +6
Search agents connect LLMs to the Internet, enabling them to access broader and more up-to-date information. However, this also introduces a new threat surface: unreliable search r…
Stop Before You Fail: Operational Capability Boundaries for Mitigating Unproductive Reasoning in Large Reasoning Models
Qingjie Zhang, Yujia Fu, Yang Wang +5
Current answering paradigms for Large Reasoning Models (LRMs) often fail to account for the fact that some questions may lie beyond the model's operational capability boundary, lea…
Taming OpenClaw: Security Analysis and Mitigation of Autonomous LLM Agent Threats
Xinhao Deng, Yixiang Zhang, Jiaqing Wu +15
Autonomous Large Language Model (LLM) agents, exemplified by OpenClaw, demonstrate remarkable capabilities in executing complex, long-horizon tasks. However, their tightly coupled…
Speculating LLMs' Chinese Training Data Pollution from Their Tokens
Qingjie Zhang, Di Wang, Haoting Qian +7
Tokens are basic elements in the datasets for LLM training. It is well-known that many tokens representing Chinese phrases in the vocabulary of GPT (4o/4o-mini/o1/o3/4.5/4.1/o4-min…
Understanding the Dilemma of Unlearning for Large Language Models
Qingjie Zhang, Haoting Qian, Zhicong Huang +5
Unlearning seeks to remove specific knowledge from large language models (LLMs), but its effectiveness remains contested. On one side, "forgotten" knowledge can often be recovered…