2 citations · 5 across the 5 of their papers we have counts for
5 papers
Generative AI Act II: Test Time Scaling Drives Cognition Engineering
Shijie Xia, Yiwei Qin, Xuefeng Li +11
The first generation of Large Language Models - what might be called "Act I" of generative AI (2020-2023) - achieved remarkable success through massive parameter and data scaling,…
ToRL: Scaling Tool-Integrated RL
Xuefeng Li, Haoyang Zou, Pengfei Liu
We introduce ToRL (Tool-Integrated Reinforcement Learning), a framework for training large language models (LLMs) to autonomously use computational tools via reinforcement learning…
LIMR: Less is More for RL Scaling
Xuefeng Li, Haoyang Zou, Pengfei Liu
In this paper, we ask: what truly determines the effectiveness of RL training data for enhancing language models' reasoning capabilities? While recent advances like o1, Deepseek R1…
PC Agent: While You Sleep, AI Works -- A Cognitive Journey into Digital World
Yanheng He, Jiahe Jin, Shijie Xia +5
Imagine a world where AI can handle your work while you sleep - organizing your research materials, drafting a report, or creating a presentation you need for tomorrow. However, wh…
O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?
Zhen Huang, Haoyang Zou, Xuefeng Li +7
This paper presents a critical examination of current approaches to replicating OpenAI's O1 model capabilities, with particular focus on the widespread but often undisclosed use of…