4 papers
Hindsight Memory-PRM: Supervising Memory Management with Auditable Hindsight Credit
Haoxuan Jia, Yang Liu, Yingguang Yang +14
Memory operations of long-horizon LLM agents are hard to supervise: an operation's value is unobservable when it is taken. But they are special -- they leave machine-readable evide…
Paying for Honesty Without Knowing the Truth: Reputation-Penalty Design for LLM Marketplace Agents
Mingdai Yang, Shicheng Fan, Kejing Yu +5
LLM agents increasingly act as autonomous merchants that write their own product listings, and under competitive pressure, they fabricate attributes to win sales. Even under instru…
GuideSkill: Evolving Executable LLM Agent Skills for Guideline-Grounded Clinical Reasoning
Lang Cao, Yuhao Shen, Tianyang Luo +3
Clinical practice guidelines (CPGs) encode diagnostic criteria, but LLM systems typically retrieve guideline text or absorb it through training rather than execute its rules. We in…
Characterizing Rhetorical Misalignment in Decision-Making with Language Models
Zirui Cheng, Joey Chan, Simo Du +3
Human decision-making is often shaped by a range of well-documented cognitive biases. As large language models (LLMs) become increasingly integrated into high-stakes human-AI decis…