4 papers
RPO: Decoupling Rollout and Inference Policies for LLM Reasoning
Jingchu Wang, Bingbing Xu, Yige Yuan +4
Existing reinforcement learning methods for LLM reasoning implicitly assume that the policy generating training trajectories should coincide with the one producing inference respon…
Towards Robust Process Reward Modeling via Noise-aware Learning
Bin Xie, Bingbing Xu, Xueyun Tian +2
Process Reward Models (PRMs) have achieved strong results in complex reasoning, but are bottlenecked by costly process-level supervision. A widely used alternative, Monte Carlo Est…
From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment
Bin Xie, Bingbing Xu, Yige Yuan +2
Inference-time alignment methods have gained significant attention for their efficiency and effectiveness in aligning large language models (LLMs) with human preferences. However,…
DHPrep: Deep Hawkes Process based Dynamic Network Representation
Ruixuan Han, Hongxiang Li, Bin Xie
Networks representation aims to encode vertices into a low-dimensional space, while preserving the original network structures and properties. Most existing methods focus on static…