Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
GraphPO: Graph-based Policy Optimization for Reasoning Models
Yuliang Zhan, Xinyu Tang, Jian Li +7
Reinforcement Learning with Verifiable Rewards (RLVR) has become a standard paradigm for enhancing the capability of large reasoning models. RLVR typically samples responses indepe…
cs.CL2025
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs
Hao Sun, Yunyi Shen, Jean-Francois Ton +1
Large Language Models (LLMs) have made substantial strides in structured tasks through Reinforcement Learning (RL), demonstrating proficiency in mathematical reasoning and code gen…
cs.CL2025
Reviving The Classics: Active Reward Modeling in Large Language Model Alignment
Yunyi Shen, Hao Sun, Jean-François Ton
Building neural reward models from human preferences is a pivotal component in reinforcement learning from human feedback (RLHF) and large language model alignment research. Given…