most citedUnsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents

1 citations · 1 across the 4 of their papers we have counts for

collaborators

12 papers

cs.CR20261 cited

Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents

Xu Li, Simon Yu, Minzhou Pan +5

LLM-based agents are becoming increasingly capable, yet their safety lags behind. This creates a gap between what agents can do and should do. This gap widens as agents engage in m…

cs.CL2026

When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents

Jaylen Jones, Zhehao Zhang, Yuting Ning +6

Although computer-use agents (CUAs) hold significant potential to automate increasingly complex OS workflows, they can demonstrate unsafe unintended behaviors that deviate from exp…

cs.CL2026

Autonomous Continual Learning for Environment Adaptation of Computer-Use Agents

Tianci Xue, Zeyi Liao, Tianneng Shi +5

Real-world digital environments are highly diverse and dynamic. These characteristics cause agents to frequently encounter unseen environments and distribution shifts, making conti…

cs.CR2026

MalTool: Malicious Tool Attacks on LLM Agents

Yuepeng Hu, Yuqi Jia, Mengyuan Li +2

In a malicious tool attack, an attacker uploads a malicious tool to a distribution platform; once a user inadvertently installs the tool and the LLM agent selects it during task ex…

cs.AI2026

Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and Opportunities

Changdae Oh, Seongheon Park, To Eun Kim +8

Uncertainty quantification (UQ) for large language models (LLMs) is a key building block for safety guardrails of daily LLM applications. Yet, even as LLM agents are increasingly d…

cs.AI2026

Signal in the Noise: Polysemantic Interference Transfers and Predicts Cross-Model Influence

Bofan Gong, Shiyang Lai, James Evans +1

Polysemanticity is pervasive in language models and remains a major challenge for interpretation and model behavioral control. Leveraging sparse autoencoders (SAEs), we map the pol…