Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Algorithmic Recourse of In-Context Learning for Tabular Data
Wenshuo Dong, Jiaming Zhang, Shaopeng Fu +3
As predictive models are increasingly deployed in high-stakes settings such as credit approval, there is a growing need for post-hoc methods that provide recourse to affected indiv…
cs.LG2026
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
Yiran Xu, Yiming Ren, Zicheng Lin +8
We identify a new dimension for enhancing rollout diversity in Group Relative Policy Optimization (GRPO) for LLMs. While GRPO relies on diverse rollouts, prevailing strategies prim…
cs.LG2025
PersRM-R1: Enhance Personalized Reward Modeling with Reinforcement Learning
Mengdi Li, Guanqiao Chen, Xufeng Zhao +3
Reward models (RMs), which are central to existing post-training methods, aim to align LLM outputs with human values by providing feedback signals during fine-tuning. However, exis…