2 papers
cs.AI2025
SPPD: Self-training with Process Preference Learning Using Dynamic Value Margin
Hao Yi, Qingyang Li, Yulan Hu +3
Recently, enhancing the numerical and logical reasoning capability of Large Language Models (LLMs) has emerged as a research hotspot. Existing methods face several limitations: inf…
cs.LG2024
TSO: Self-Training with Scaled Preference Optimization
Kaihui Chen, Hao Yi, Qingyang Li +4
Enhancing the conformity of large language models (LLMs) to human preferences remains an ongoing research challenge. Recently, offline approaches such as Direct Preference Optimiza…