4 papers
O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?
Zhen Huang, Haoyang Zou, Xuefeng Li +7
This paper presents a critical examination of current approaches to replicating OpenAI's O1 model capabilities, with particular focus on the widespread but often undisclosed use of…
O1 Replication Journey: A Strategic Progress Report -- Part 1
Yiwei Qin, Xuefeng Li, Haoyang Zou +8
This paper introduces a pioneering approach to artificial intelligence research, embodied in our O1 Replication Journey. In response to the announcement of OpenAI's groundbreaking…
LLMCRIT: Teaching Large Language Models to Use Criteria
Weizhe Yuan, Pengfei Liu, Matthias Gallé
Humans follow criteria when they execute tasks, and these criteria are directly used to assess the quality of task completion. Therefore, having models learn to use criteria to pro…
The Critique of Critique
Shichao Sun, Junlong Li, Weizhe Yuan +3
Critique, as a natural language description for assessing the quality of model-generated content, has played a vital role in the training, evaluation, and refinement of LLMs. Howev…