Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
GAGPO: Generalized Advantage Grouped Policy Optimization
Siyuan Zhu, Chao Yu, Rongxin Yang +4
Reinforcement learning has become a powerful paradigm for post-training large language model agents, yet credit assignment in multi-turn environments remains a challenge. Agents of…
cs.CL2026
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
Zeyu Tang, Sang T. Truong, Deonna Owens +4
LLM fairness should be evaluated through in-situ behavioral pattern rather than standardized-test Q&A benchmarks. We show that the standardized-test paradigm can be structurally un…