works on

From the 2 of 13 linked papers with an AI index.

most citedSafety at Scale: A Comprehensive Survey of Large Model and Agent Safety

1 citations · 1 across the 9 of their papers we have counts for

collaborators
Showing cs.AIShow all

5 papers · 1 filter

cs.AI2026

Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents

Xutao Mao, Liangjie Zhao, Xiang Zheng +1

Self-improving LLM agents convert successful trajectories into persistent cross-task state. An unsafe success can thereby become reusable policy after its triggering input disappea…

cs.AI2026

Agents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Stateful Personal Agents

Xutao Mao, Liangjie Zhao, Leyao Wang +6

The paper defines persistent sycophancy, where personal agents store user‑provided claims in long‑term memory and later repeat them, and introduces the Personal Agent Sycophancy Be…

cs.AI2026

Just Ask: Curious Code Agents Reveal System Prompts in Frontier LLMs

Xiang Zheng, Yutao Wu, Hanxun Huang +5

Autonomous code agents built on large language models are reshaping software and AI development through tool use, long-horizon reasoning, and self-directed interaction. However, th…

cs.AI2026

What Happens Inside Agent Memory? Circuit Analysis from Emergence to Diagnosis

Xutao Mao, Jinman Zhao, Gerald Penn +1

Agent memory failures are silent: an LLM-based agent can produce a fluent response even when it fails to extract, retain, or retrieve the information needed across sessions. The wr…

cs.AI2025

CALM: Curiosity-Driven Auditing for Large Language Models

Xiang Zheng, Longxiang Wang, Yi Liu +3

Auditing Large Language Models (LLMs) is a crucial and challenging task. In this study, we focus on auditing black-box LLMs without access to their parameters, only to the provided…