collaborators

14 papers

cs.CR2026

Understanding and Evaluating Claw-like Agent Security Through a Computer-Systems Lens

Peizhi Niu, Wenjie Qu, Shangding Gu +14

Claw-like AI agents (e.g., OpenClaw) are always-on processes with persistent access to credentials, files, tools, and external services. They take on system-level responsibilities…

cs.AI2026

AION: Next-Generation Tasks and Practical Harness for Time Series

Tianxiang Zhan, Xiaobao Song, Tong Guan +2

Time series research is moving beyond fixed forecasting benchmarks toward realistic tasks that combine prediction, contextual reasoning, tool use, and structured decision support.…

cs.CR2026

Memory-Induced Tool-Drift in LLM Agents

Mahavir Dabas, Jihyun Jeong, Ming Jin +1

Modern LLM agents combine long-term memory for personalization with tool-calling interfaces for taking actions in the world -- a combination underpinning contemporary production sy…

cs.CL2026

MemGym: a Long-Horizon Memory Environment for LLM Agents

Wujiang Xu, Yu Wang, Kai Mei +8

Memory is a central capability for LLM agents operating across long-horizon tasks. Existing memory benchmarks predominantly evaluate retention of personalized information in multi-…

cs.AI2026

Remembering More, Risking More: Longitudinal Safety Risks in Memory-Equipped LLM Agents

Ahmad Al-Tawaha, Shangding Gu, Peizhi Niu +2

Safety evaluations of memory-equipped LLM agents typically measure within-task safety: whether an agent completes a single scenario safely, often under adversarial conditions such…

cs.LG2026

LLMs Should Express Uncertainty Explicitly

Junyu Guo, Shangding Gu, Ming Jin +2

Large language models (LLMs) often produce confident yet incorrect answers, which can lead to risky failures in real-world applications. We study whether post-training can make a m…