activity
20242026
most citedContextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards

1 citations · 3 across the 24 of their papers we have counts for

collaborators
Showing cs.CLShow all

13 papers · 1 filter

cs.CL2026

HiDiffTIR: Hierarchical Difficulty-Aware Policy Optimization for Multi-Turn Tool-Integrated Reasoning

Yucan Guo, Xiaohan Wang, Miao Su +8

Tool-Integrated Reasoning (TIR) is a fundamental capability for LLM agents to solve complex tasks by interacting with external tools iteratively. Reinforcement Learning (RL) has be…

cs.CL2026

UniMem: Complementary Episodic-to-Parametric Memory for Boundary-Agnostic Task Streams

Siyu Xia, Chenheng Zhang, Yanting Wu +8

Memory is essential for LLM agents to accumulate task experience and reuse task-specific execution strategies. However, real-world deployment over boundary-agnostic and evolving ta…

cs.CL2026

Are Full Rollouts Necessary for On-Policy Distillation?

Yaocheng Zhang, Jiajun Chai, Yuqian Fu +7

On-policy distillation (OPD) provides dense teacher feedback along student-generated rollouts rather than fixed teacher traces and has emerged as a promising post-training paradigm…

cs.CL2026

Terminal-World: Scaling Terminal-Agent Environments via Agent Skills

Zihao Cheng, Hongru Wang, Zeming Liu +6

Terminal agents extend Large Language Models with the ability to execute tasks directly in command-line environments, but their progress is bottlenecked by the scarcity of high-qua…

cs.CL2026

Implicit Hierarchical GRPO: Decoupling Tool Invocation from Execution for Tool-Integrated Mathematical Reasoning

Li Wang, Xiaohan Wang, Xiaodong Lu +5

Large language models (LLMs) have increasingly leveraged tool invocation to enhance their reasoning capabilities. However, existing approaches typically tightly couple tool invocat…

cs.CL2026

MemEvolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience Distillation

Zihao Cheng, Zeming Liu, Yingyu Shan +7

While large language model--powered agents can self-evolve by accumulating experience or by dynamically creating new assets (i.e., tools or expert agents), existing frameworks typi…