4 papers
From Profiling to Synthesis: Benchmarking Implicit Behavioral Alignment in Personalized LLM Agents
Jiajia Song, Bobo Li, Haiwen Yi +6
Large Language Models have enabled increasingly capable autonomous agents, yet personalization remains critical for making such agents practically useful. Recent benchmarks have be…
Learning to Control LLM Agent Harnesses with Offline Reinforcement Learning
Haiwen Yi, Xinyuan Song
Large language model (LLM) agents are usually improved by changing prompts, models, or hand-written workflows, while the execution harness around the model is treated as fixed infr…
ManifoldFlow: SPD-Relaxed Stiefel Layers with Learnable Singular Spectrum
Haiwen Yi, Xinyuan Song
Orthogonal and Stiefel layers give neural weights exact spectral control, but they also impose a strong modeling constraint: all represented singular values are fixed at one. Many…
Measuring Harness-Induced Belief Divergence in Multi-Step LLM Agents
Haiwen Yi, Xinyuan Song
Software-agent benchmarks usually report whether an agent solves a task, but the agent reaches that outcome through a harness that controls what it sees, which actions it can take,…