1 citations · 1 across the 4 of their papers we have counts for
Showing 2026 · cs.CLShow all
2 papers · 2 filters
cs.CL2026
GAGPO: Generalized Advantage Grouped Policy Optimization
Siyuan Zhu, Chao Yu, Rongxin Yang +4
Reinforcement learning has become a powerful paradigm for post-training large language model agents, yet credit assignment in multi-turn environments remains a challenge. Agents of…
cs.CL2026
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
Zeyu Tang, Sang T. Truong, Deonna Owens +4
LLM fairness should be evaluated through in-situ behavioral pattern rather than standardized-test Q&A benchmarks. We show that the standardized-test paradigm can be structurally un…