9 papers
The 2026 Singapore Consensus on Global AI Safety Research Priorities
Stephen Casper, Oskar Galeev, Yoshua Bengio +117
Frontier AI capabilities and autonomy are advancing rapidly. A growing number of real-world incidents make a trusted AI ecosystem essential to embracing AI with confidence. The 202…
Offline Safe Policy Optimization From Heterogeneous Feedback
Ze Gong, Pradeep Varakantham, Akshat Kumar
Offline Preference-based Reinforcement Learning (PbRL) learns rewards and policies aligned with human preferences without the need for extensive reward engineering and direct inter…
FairVizARD: A Visualization System for Assessing Multi-Party Fairness of Ride-Sharing Matching Algorithms
Ashwin Kumar, Sanket Shah, Meghna Lowalekar +3
There is growing interest in algorithms that match passengers with drivers in ride-sharing problems and their fairness for the different parties involved (passengers, drivers, and…
On Minimizing Adversarial Counterfactual Error in Adversarial RL
Roman Belaire, Arunesh Sinha, Pradeep Varakantham
Deep Reinforcement Learning (DRL) policies are highly susceptible to adversarial noise in observations, which poses significant risks in safety-critical scenarios. The challenge in…
Offline Safe Reinforcement Learning Using Trajectory Classification
Ze Gong, Akshat Kumar, Pradeep Varakantham
Offline safe reinforcement learning (RL) has emerged as a promising approach for learning safe behaviors without engaging in risky online interactions with the environment. Most ex…
Bootstrapping Language Models with DPO Implicit Rewards
Changyu Chen, Zichen Liu, Chao Du +5
Human alignment in large language models (LLMs) is an active area of research. A recent groundbreaking work, direct preference optimization (DPO), has greatly simplified the proces…