3 papers
cs.AI2026
Harness-agnostic detection and immunization of reward hacking in self-evolving language models
Rongxin Yang, Yang Liu, Shang Luo +10
Self-evolving language models improve by proposing candidate updates and keeping whatever raises a visible score. When that score is an imperfect proxy for the capability one actua…
cs.CR2026
Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents
Chenhao Wu, Haoxuan Jia, Yang Liu +11
Large language model agents are increasingly deployed as autonomous loops. Starting from one human goal, such a system repeatedly discovers work, plans, executes tool calls, verifi…
cs.CL2026
Hindsight Memory-PRM: Supervising Memory Management with Auditable Hindsight Credit
Haoxuan Jia, Yang Liu, Yingguang Yang +14
Memory operations of long-horizon LLM agents are hard to supervise: an operation's value is unobservable when it is taken. But they are special -- they leave machine-readable evide…