4 papers
ANCHOR: Automated Alignment Auditing for CLI Agents on Real-World Harm
Kefan Song, Yanjun Qi
Autonomous CLI agents can now execute hundreds of actions across multi-hour sessions: writing code, executing shell commands, browsing the web, and managing cloud infrastructure, a…
Reward Is Enough: LLMs Are In-Context Reinforcement Learners
Kefan Song, Amir Moeini, Peng Wang +4
Reinforcement learning (RL) is a framework for solving sequential decision-making problems. In this work, we demonstrate that, surprisingly, RL emerges during the inference time of…
Group Fairness in Multi-Task Reinforcement Learning
Kefan Song, Runnan Jiang, Rohan Chandra +1
This paper addresses a critical societal consideration in the application of Reinforcement Learning (RL): ensuring equitable outcomes across different demographic groups in multi-t…
Towards Large Language Models that Benefit for All: Benchmarking Group Fairness in Reward Models
Kefan Song, Jin Yao, Runnan Jiang +2
As Large Language Models (LLMs) become increasingly powerful and accessible to human users, ensuring fairness across diverse demographic groups, i.e., group fairness, is a critical…