1 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.LG2025★ 1 cited
Process Reward Models for LLM Agents: Practical Framework and Directions
Sanjiban Choudhury
We introduce Agent Process Reward Models (AgentPRM), a simple and scalable framework for training LLM agents to continually improve through interactions. AgentPRM follows a lightwe…
cs.LG2025
Aligning LLMs with Domain Invariant Reward Models
David Wu, Sanjiban Choudhury
Aligning large language models (LLMs) to human preferences is challenging in domains where preference data is unavailable. We address the problem of learning reward models for such…
cs.LG2024★ 1 cited
Better than Your Teacher: LLM Agents that learn from Privileged AI Feedback
Sanjiban Choudhury, Paloma Sodhi
While large language models (LLMs) show impressive decision-making abilities, current methods lack a mechanism for automatic self-improvement from errors during task execution. We…