4 papers
GAGPO: Generalized Advantage Grouped Policy Optimization
Siyuan Zhu, Chao Yu, Rongxin Yang +4
Reinforcement learning has become a powerful paradigm for post-training large language model agents, yet credit assignment in multi-turn environments remains a challenge. Agents of…
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
Zeyu Tang, Sang T. Truong, Deonna Owens +4
LLM fairness should be evaluated through in-situ behavioral pattern rather than standardized-test Q&A benchmarks. We show that the standardized-test paradigm can be structurally un…
A Framework for Objective-Driven Dynamical Stochastic Fields
Yibo Jacky Zhang, Sanmi Koyejo
Fields offer a versatile approach for describing complex systems composed of interacting and dynamic components. In particular, some of these dynamical and stochastic systems may e…
Probing Human Visual Robustness with Neurally-Guided Deep Neural Networks
Zhenan Shao, Linjian Ma, Yiqing Zhou +4
Humans effortlessly navigate the dynamic visual world, yet deep neural networks (DNNs), despite excelling at many visual tasks, are surprisingly vulnerable to minor image perturbat…