6 papers
Learn More with Less: Uncertainty Consistency Guided Query Selection for RLVR
Hao Yi, Yulan Hu, Xin Li +3
Large Language Models (LLMs) have recently improved mathematical reasoning through Reinforcement Learning with Verifiable Reward (RLVR). However, existing RLVR algorithms require l…
Coarse-to-Fine Process Reward Modeling for Mathematical Reasoning
Yulan Hu, Sheng Ouyang, Jinman Zhao +1
The Process Reward Model (PRM) plays a crucial role in mathematical reasoning tasks, requiring high-quality supervised process data. However, we observe that reasoning steps genera…
Pieceformer: Similarity-Driven Knowledge Transfer via Scalable Graph Transformer in VLSI
Hang Yang, Yusheng Hu, Yong Liu +2
Accurate graph similarity is critical for knowledge transfer in VLSI design, enabling the reuse of prior solutions to reduce engineering effort and turnaround time. We propose Piec…
Towards Reward Fairness in RLHF: From a Resource Allocation Perspective
Sheng Ouyang, Yulan Hu, Ge Chen +3
Rewards serve as proxies for human preferences and play a crucial role in Reinforcement Learning from Human Feedback (RLHF). However, if these rewards are inherently imperfect, exh…
Perfect Alignment May be Poisonous to Graph Contrastive Learning
Jingyu Liu, Huayi Tang, Yong Liu
Graph Contrastive Learning (GCL) aims to learn node representations by aligning positive pairs and separating negative ones. However, few of researchers have focused on the inner l…
GUNDAM: Aligning Large Language Models with Graph Understanding
Sheng Ouyang, Yulan Hu, Ge Chen +1
Large Language Models (LLMs) have achieved impressive results in processing text data, which has sparked interest in applying these models beyond textual data, such as graphs. In t…