Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
From Profiling to Synthesis: Benchmarking Implicit Behavioral Alignment in Personalized LLM Agents
Jiajia Song, Bobo Li, Haiwen Yi +6
Large Language Models have enabled increasingly capable autonomous agents, yet personalization remains critical for making such agents practically useful. Recent benchmarks have be…
cs.AI2026
Measuring Harness-Induced Belief Divergence in Multi-Step LLM Agents
Haiwen Yi, Xinyuan Song
Software-agent benchmarks usually report whether an agent solves a task, but the agent reaches that outcome through a harness that controls what it sees, which actions it can take,…