11 citations · 43 across the 60 of their papers we have counts for
13 papers · 1 filter
Self-Abstraction from Grounded Experience for Plan-Guided Policy Refinement
Hiroaki Hayashi, Bo Pang, Wenting Zhao +6
Large language model (LLM) based agents are increasingly used to tackle software engineering tasks that require multi-step reasoning and code modification, demonstrating promising…
Reasoning Curriculum: Bootstrapping Broad LLM Reasoning from Math
Bo Pang, Deqian Kong, Silvio Savarese +2
Reinforcement learning (RL) can elicit strong reasoning in large language models (LLMs), yet most open efforts focus on math and code. We propose Reasoning Curriculum, a simple two…
SCUBA: Salesforce Computer Use Benchmark
Yutong Dai, Krithika Ramakrishnan, Jing Gu +8
We introduce SCUBA, a benchmark designed to evaluate computer-use agents on customer relationship management (CRM) workflows within the Salesforce platform. SCUBA contains 300 task…
UserRL: Training Interactive User-Centric Agent via Reinforcement Learning
Cheng Qian, Zuxin Liu, Akshara Prabhakar +10
Reinforcement learning (RL) has shown promise in training agentic models that move beyond static benchmarks to engage in dynamic, multi-turn interactions. Yet, the ultimate value o…
MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models
Zhiwei Liu, Jielin Qiu, Shiyu Wang +9
The rapid rise of Large Language Models (LLMs)-based intelligent agents underscores the need for robust, scalable evaluation frameworks. Existing methods rely on static benchmarks…
UserBench: An Interactive Gym Environment for User-Centric Agents
Cheng Qian, Zuxin Liu, Akshara Prabhakar +9
Large Language Models (LLMs)-based agents have made impressive progress in reasoning and tool use, enabling them to solve complex tasks. However, their ability to proactively colla…