From the 2 of 13 linked papers with an AI index.
1 citations · 1 across the 9 of their papers we have counts for
5 papers · 1 filter
Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents
Xutao Mao, Liangjie Zhao, Xiang Zheng +1
Self-improving LLM agents convert successful trajectories into persistent cross-task state. An unsafe success can thereby become reusable policy after its triggering input disappea…
Agents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Stateful Personal Agents
Xutao Mao, Liangjie Zhao, Leyao Wang +6
The paper defines persistent sycophancy, where personal agents store user‑provided claims in long‑term memory and later repeat them, and introduces the Personal Agent Sycophancy Be…
Just Ask: Curious Code Agents Reveal System Prompts in Frontier LLMs
Xiang Zheng, Yutao Wu, Hanxun Huang +5
Autonomous code agents built on large language models are reshaping software and AI development through tool use, long-horizon reasoning, and self-directed interaction. However, th…
What Happens Inside Agent Memory? Circuit Analysis from Emergence to Diagnosis
Xutao Mao, Jinman Zhao, Gerald Penn +1
Agent memory failures are silent: an LLM-based agent can produce a fluent response even when it fails to extract, retain, or retrieve the information needed across sessions. The wr…
CALM: Curiosity-Driven Auditing for Large Language Models
Xiang Zheng, Longxiang Wang, Yi Liu +3
Auditing Large Language Models (LLMs) is a crucial and challenging task. In this study, we focus on auditing black-box LLMs without access to their parameters, only to the provided…