Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training
Ashish Jain, Armaan Sandhu
Interactive agent benchmarks and multi-turn reinforcement learning increasingly place a second language model in the role of the user. This simulated user controls what information…
cs.AI2026
Measuring Reward Hacking and Reasoning-Answer Decoupling Under Position-Confounded Optimization
Suyash Maniyar, Armaan Sandhu, Abhishek Mishra
When a reward is correct on every training example yet consistent with more than one goal, a model can acquire an unintended one, a failure known as goal misgeneralization. Endpoin…