From the 1 of 4 linked papers with an AI index.
4 papers
Algorithmic Recourse of In-Context Learning for Tabular Data
Wenshuo Dong, Jiaming Zhang, Shaopeng Fu +3
The paper introduces a theoretical and practical framework for providing algorithmic recourse on tabular data using in-context learning with large language models, proposing a zero…
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
Yiming Ren, Yiran Xu, Zicheng Lin +8
We identify a new dimension for enhancing rollout diversity in Group Relative Policy Optimization (GRPO) for LLMs. While GRPO relies on diverse rollouts, prevailing strategies prim…
Curriculum-RLAIF: Curriculum Alignment with Reinforcement Learning from AI Feedback
Jiaye Lin, Mengdi Li, Xufeng Zhao +4
Reward models trained through Reinforcement Learning from AI Feedback (RLAIF) methods frequently suffer from limited generalizability, which hinders the alignment performance of po…
PersRM-R1: Enhance Personalized Reward Modeling with Reinforcement Learning
Mengdi Li, Guanqiao Chen, Xufeng Zhao +3
Reward models (RMs), which are central to existing post-training methods, aim to align LLM outputs with human values by providing feedback signals during fine-tuning. However, exis…