activity
20242026
most citedYour Group-Relative Advantage Is Biased

1 citations · 4 across the 43 of their papers we have counts for

collaborators
Showing 2026Show all

27 papers · 1 filter

cs.SE2026

Tool Retrievers Are Underestimated: Annotation Expansion Reveals True Capability

Yanyu Zhu, Chenheng Zhang, Shaoshen Chen +8

In open-world scenarios with massive and evolving tool repositories, tool-augmented large language models rely on a retriever to surface relevant tools for a given query. Because s…

cs.LG2026

One Step, One Lead: Mitigating Higher-Order Interference in Multi-Domain Reinforcement Learning via Cross-Step Control

Zihan Lin, Xiaohan Wang, Jie Cao +4

Reinforcement learning (RL) across multiple domains can broaden the reasoning capabilities of large language models (LLMs), yet joint training often degrades individual-domain perf…

cs.CL2026

HiDiffTIR: Hierarchical Difficulty-Aware Policy Optimization for Multi-Turn Tool-Integrated Reasoning

Yucan Guo, Xiaohan Wang, Miao Su +8

Tool-Integrated Reasoning (TIR) is a fundamental capability for LLM agents to solve complex tasks by interacting with external tools iteratively. Reinforcement Learning (RL) has be…

cs.AI2026

ATLAS: Dual-Horizon Diagnostic Evaluation for Industrial Tool-Use Agents

Wei Chen, Peilun Zhou, Zhaoyu Hu +8

Large language model (LLM) agents are increasingly deployed in user-facing services that require iterative tool use under dynamic business conditions. Reliable evaluation is essent…

cs.CL2026

Behavior2Trip: Towards Personalized Travel Planning via User Behavior Trajectory

Zihao Cheng, Yingyu Shan, Hongru Wang +6

Travel planning agents assist users in generating personalized travel plans by modeling their individual preferences. Existing agents either rely on explicit user instructions or e…

cs.CL2026

When Not to Imitate: Boundary-Aware Skill Memory for Reliable Tool-Use LLM Agents

Zihan Lin, Zhenyu Chen, Jiawen Wei +6

Extracting skills from past successes is critical for the efficient evolution of Large Language Model (LLM) agents. Prevailing agent self-evolution paradigms typically rely on a co…