3 papers
cs.LG2025
Indeterminate Probability Theory
Tao Yang, Chuang Liu, Xiaofeng Ma +8
Complex continuous or mixed joint distributions (e.g., P(Y | z_1, z_2, ..., z_N)) generally lack closed-form solutions, often necessitating approximations such as MCMC. This paper…
cs.CL2025
FreePRM: Training Process Reward Models Without Ground Truth Process Labels
Lin Sun, Chuang Liu, Xiaofeng Ma +3
Recent advancements in Large Language Models (LLMs) have demonstrated that Process Reward Models (PRMs) play a crucial role in enhancing model performance. However, training PRMs t…
cs.CL2025
BPO: Revisiting Preference Modeling in Direct Preference Optimization
Lin Sun, Chuang Liu, Peng Liu +3
Direct Preference Optimization (DPO) have emerged as a popular method for aligning Large Language Models (LLMs) with human preferences. While DPO effectively preserves the relative…