4 papers
Algorithmic Recourse of In-Context Learning for Tabular Data
Wenshuo Dong, Jiaming Zhang, Shaopeng Fu +3
As predictive models are increasingly deployed in high-stakes settings such as credit approval, there is a growing need for post-hoc methods that provide recourse to affected indiv…
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
Yiran Xu, Yiming Ren, Zicheng Lin +8
We identify a new dimension for enhancing rollout diversity in Group Relative Policy Optimization (GRPO) for LLMs. While GRPO relies on diverse rollouts, prevailing strategies prim…
PersRM-R1: Enhance Personalized Reward Modeling with Reinforcement Learning
Mengdi Li, Guanqiao Chen, Xufeng Zhao +3
Reward models (RMs), which are central to existing post-training methods, aim to align LLM outputs with human values by providing feedback signals during fine-tuning. However, exis…
Curriculum-RLAIF: Curriculum Alignment with Reinforcement Learning from AI Feedback
Jiaye Lin, Mengdi Li, Xufeng Zhao +4
Reward models trained through Reinforcement Learning from AI Feedback (RLAIF) methods frequently suffer from limited generalizability, which hinders the alignment performance of po…