4 papers · 1 filter
Beyond Rejection Sampling: Trajectory Fusion for Scaling Mathematical Reasoning
Jie Deng, Hanshuang Tong, Jun Li +4
Large language models (LLMs) have made impressive strides in mathematical reasoning, often fine-tuned using rejection sampling that retains only correct reasoning trajectories. Whi…
FreePRM: Training Process Reward Models Without Ground Truth Process Labels
Lin Sun, Chuang Liu, Xiaofeng Ma +3
Recent advancements in Large Language Models (LLMs) have demonstrated that Process Reward Models (PRMs) play a crucial role in enhancing model performance. However, training PRMs t…
BPO: Revisiting Preference Modeling in Direct Preference Optimization
Lin Sun, Chuang Liu, Peng Liu +3
Direct Preference Optimization (DPO) have emerged as a popular method for aligning Large Language Models (LLMs) with human preferences. While DPO effectively preserves the relative…
Seed1.5-Thinking: Advancing Superb Reasoning Models with Reinforcement Learning
ByteDance Seed, :, Jiaze Chen +267
We introduce Seed1.5-Thinking, capable of reasoning through thinking before responding, resulting in improved performance on a wide range of benchmarks. Seed1.5-Thinking achieves 8…