3 papers
cs.AI2026
UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training
Ashish Jain, Armaan Sandhu
Interactive agent benchmarks and multi-turn reinforcement learning increasingly place a second language model in the role of the user. This simulated user controls what information…
cs.CV2026
Targeting the Attention Heads Behind Object Hallucination in LLaVA
Armaan Sandhu, Abhilasha Senapati, Hima Kammachi
Vision-language models such as LLaVA-1.5-7B often hallucinate objects absent from the image when generating captions. We ask whether an interpretability diagnosis of this failure c…
cs.AI2026
Measuring Reward Hacking and Reasoning-Answer Decoupling Under Position-Confounded Optimization
Suyash Maniyar, Armaan Sandhu, Abhishek Mishra
When a reward is correct on every training example yet consistent with more than one goal, a model can acquire an unintended one, a failure known as goal misgeneralization. Endpoin…