14 papers
Understanding and Evaluating Claw-like Agent Security Through a Computer-Systems Lens
Peizhi Niu, Wenjie Qu, Shangding Gu +14
Claw-like AI agents (e.g., OpenClaw) are always-on processes with persistent access to credentials, files, tools, and external services. They take on system-level responsibilities…
AION: Next-Generation Tasks and Practical Harness for Time Series
Tianxiang Zhan, Xiaobao Song, Tong Guan +2
Time series research is moving beyond fixed forecasting benchmarks toward realistic tasks that combine prediction, contextual reasoning, tool use, and structured decision support.…
Memory-Induced Tool-Drift in LLM Agents
Mahavir Dabas, Jihyun Jeong, Ming Jin +1
Modern LLM agents combine long-term memory for personalization with tool-calling interfaces for taking actions in the world -- a combination underpinning contemporary production sy…
MemGym: a Long-Horizon Memory Environment for LLM Agents
Wujiang Xu, Yu Wang, Kai Mei +8
Memory is a central capability for LLM agents operating across long-horizon tasks. Existing memory benchmarks predominantly evaluate retention of personalized information in multi-…
Remembering More, Risking More: Longitudinal Safety Risks in Memory-Equipped LLM Agents
Ahmad Al-Tawaha, Shangding Gu, Peizhi Niu +2
Safety evaluations of memory-equipped LLM agents typically measure within-task safety: whether an agent completes a single scenario safely, often under adversarial conditions such…
LLMs Should Express Uncertainty Explicitly
Junyu Guo, Shangding Gu, Ming Jin +2
Large language models (LLMs) often produce confident yet incorrect answers, which can lead to risky failures in real-world applications. We study whether post-training can make a m…