most citedToRL: Scaling Tool-Integrated RL

2 citations · 5 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CL2025

Generative AI Act II: Test Time Scaling Drives Cognition Engineering

Shijie Xia, Yiwei Qin, Xuefeng Li +11

The first generation of Large Language Models - what might be called "Act I" of generative AI (2020-2023) - achieved remarkable success through massive parameter and data scaling,…

cs.CL20252 cited

ToRL: Scaling Tool-Integrated RL

Xuefeng Li, Haoyang Zou, Pengfei Liu

We introduce ToRL (Tool-Integrated Reinforcement Learning), a framework for training large language models (LLMs) to autonomously use computational tools via reinforcement learning…

cs.LG20251 cited

LIMR: Less is More for RL Scaling

Xuefeng Li, Haoyang Zou, Pengfei Liu

In this paper, we ask: what truly determines the effectiveness of RL training data for enhancing language models' reasoning capabilities? While recent advances like o1, Deepseek R1…

cs.AI20241 cited

PC Agent: While You Sleep, AI Works -- A Cognitive Journey into Digital World

Yanheng He, Jiahe Jin, Shijie Xia +5

Imagine a world where AI can handle your work while you sleep - organizing your research materials, drafting a report, or creating a presentation you need for tomorrow. However, wh…

cs.CL20241 cited

O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?

Zhen Huang, Haoyang Zou, Xuefeng Li +7

This paper presents a critical examination of current approaches to replicating OpenAI's O1 model capabilities, with particular focus on the widespread but often undisclosed use of…