From the 2 of 10 linked papers with an AI index.
10 papers
SAFETY SENTRY: Context-Aware Human Intervention via EXECUTE-ASK-REFUSE Routing
Tianyu Chen, Chujia Hu, Wenjie Wang
The paper introduces Safety Sentry, a lightweight guard model for large language model agents that decides per action whether to execute, ask the user, or refuse, using a single de…
MemoHarness: Agent Harnesses That Learn from Experience
Yue Huang, Wenjie Wang, Han Bao +7
MemoHarness is a framework that automatically adapts the control layer (harness) of large language model agents by learning from past executions, using a dual‑layer experience bank…
From Passive Retrieval to Active Memory Navigation: Learning to Use Memory as a Structured Action Space
Yue Xu, Yutao Sun, Yihao Liu +7
Long-term user memory is essential for personalized conversational agents, yet many memory systems still expose memory through passive retrieval interfaces, making the model a cons…
AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security
Dongrui Liu, Yu Li, Zhonghao Yang +47
Modern open-world agents such as OpenClaw exhibit powerful cross-environment execution capabilities yet introduce broad new safety risk sources. Meanwhile, advanced frontier AI mod…
Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents
Yongxiang Li, Moxin Li, Zhixin Ma +4
Large Language Model (LLM) agents remain vulnerable to safety threats from the external environment, where attackers inject adversarial content into external observations such as t…
AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security
Dongrui Liu, Qihan Ren, Chen Qian +40
The rise of AI agents introduces complex safety and security challenges arising from autonomous tool use and environmental interactions. Current guardrail models lack agentic risk…