most citedO1 Replication Journey: A Strategic Progress Report -- Part 1

3 citations · 8 across the 10 of their papers we have counts for

collaborators

12 papers

cs.AI2025

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Haoming Wang, Haoyang Zou, Huatong Song +109

The development of autonomous agents for graphical user interfaces (GUIs) presents major challenges in artificial intelligence. While recent advances in native agent models have sh…

cs.CL2025

Interaction as Intelligence: Deep Research With Human-AI Partnership

Lyumanshan Ye, Xiaojie Cai, Xinkai Wang +23

This paper introduces "Interaction as Intelligence" research series, presenting a reconceptualization of human-AI relationships in deep research tasks. Traditional approaches treat…

cs.CL2025

CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning

Congmin Zheng, Jiachen Zhu, Jianghao Lin +6

Process Reward Models (PRMs) play a central role in evaluating and guiding multi-step reasoning in large language models (LLMs), especially for mathematical problem solving. Howeve…

cs.CL2025

Generative AI Act II: Test Time Scaling Drives Cognition Engineering

Shijie Xia, Yiwei Qin, Xuefeng Li +11

The first generation of Large Language Models - what might be called "Act I" of generative AI (2020-2023) - achieved remarkable success through massive parameter and data scaling,…

cs.CL20252 cited

ToRL: Scaling Tool-Integrated RL

Xuefeng Li, Haoyang Zou, Pengfei Liu

We introduce ToRL (Tool-Integrated Reinforcement Learning), a framework for training large language models (LLMs) to autonomously use computational tools via reinforcement learning…

cs.LG20251 cited

LIMR: Less is More for RL Scaling

Xuefeng Li, Haoyang Zou, Pengfei Liu

In this paper, we ask: what truly determines the effectiveness of RL training data for enhancing language models' reasoning capabilities? While recent advances like o1, Deepseek R1…