5 papers
Learn More with Less: Uncertainty Consistency Guided Query Selection for RLVR
Hao Yi, Yulan Hu, Xin Li +3
Large Language Models (LLMs) have recently improved mathematical reasoning through Reinforcement Learning with Verifiable Reward (RLVR). However, existing RLVR algorithms require l…
Pieceformer: Similarity-Driven Knowledge Transfer via Scalable Graph Transformer in VLSI
Hang Yang, Yusheng Hu, Yong Liu +2
Accurate graph similarity is critical for knowledge transfer in VLSI design, enabling the reuse of prior solutions to reduce engineering effort and turnaround time. We propose Piec…
Towards Reward Fairness in RLHF: From a Resource Allocation Perspective
Sheng Ouyang, Yulan Hu, Ge Chen +3
Rewards serve as proxies for human preferences and play a crucial role in Reinforcement Learning from Human Feedback (RLHF). However, if these rewards are inherently imperfect, exh…
Coarse-to-Fine Process Reward Modeling for Mathematical Reasoning
Yulan Hu, Sheng Ouyang, Jinman Zhao +1
The Process Reward Model (PRM) plays a crucial role in mathematical reasoning tasks, requiring high-quality supervised process data. However, we observe that reasoning steps genera…
TSO: Self-Training with Scaled Preference Optimization
Kaihui Chen, Hao Yi, Qingyang Li +4
Enhancing the conformity of large language models (LLMs) to human preferences remains an ongoing research challenge. Recently, offline approaches such as Direct Preference Optimiza…