4 papers
Harness-agnostic detection and immunization of reward hacking in self-evolving language models
Rongxin Yang, Yang Liu, Shang Luo +10
Self-evolving language models improve by proposing candidate updates and keeping whatever raises a visible score. When that score is an imperfect proxy for the capability one actua…
Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents
Chenhao Wu, Haoxuan Jia, Yang Liu +11
Large language model agents are increasingly deployed as autonomous loops. Starting from one human goal, such a system repeatedly discovers work, plans, executes tool calls, verifi…
Hindsight Memory-PRM: Supervising Memory Management with Auditable Hindsight Credit
Haoxuan Jia, Yang Liu, Yingguang Yang +14
Memory operations of long-horizon LLM agents are hard to supervise: an operation's value is unobservable when it is taken. But they are special -- they leave machine-readable evide…
ExTax: Explainable Disinformation Detection via Persuasion, Emotion, and Narrative Role Taxonomies
Shang Luo, Yingguang Yang, Zhenchen Sun +8
The democratization of LLMs has accelerated the generation and circulation of highly fluent disinformation, making traditional syntax-semantic verification increasingly insufficien…