Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Efficient Reinforcement Learning for Long-Horizon Tool-Use Agentic Tasks
Zelei Cheng, Amritansh Mishra, Sambit Sahu +1
Long-horizon tool-using agents must reason over user goals, domain policies, tool calls, simulator state, and delayed verifiable rewards. Reinforcement learning (RL) is a natural f…
cs.LG2026
RECAP: Regression Evaluation for Continual Adaptation of Prompts
Harsh Deshpande, Kushal Chawla, Sangwoo Cho +2
Production agentic systems routinely face evolving constraints and must comply from the very next interaction. Scenarios like a tool-call notification changing a compliance thresho…