activity
20182026
most citedWhen Humans Aren't Optimal: Robots that Collaborate with Risk-Aware Humans

57 citations · 448 across the 99 of their papers we have counts for

collaborators
Showing cs.AIShow all

13 papers · 1 filter

cs.AI2026

SPIRAL: Learning to Search and Aggregate

Jubayer Ibn Hamid, Ifdita Hasan Orney, Michael Y. Li +5

Language model reasoning can be substantially improved at test time via scaffolds that scale inference compute across different primitives -- sequential reasoning within a trace, i…

cs.AI2026

Poly-EPO: Training Exploratory Reasoning Models

Ifdita Hasan Orney, Jubayer Ibn Hamid, Shreya S Ramanujam +5

Exploration is a cornerstone of learning from experience: it enables agents to find solutions to complex problems, generalize to novel ones, and scale performance with test-time co…

cs.AI20251 cited

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning

Bidipta Sarkar, Warren Xia, C. Karen Liu +1

Communicating in natural language is a powerful tool in multi-agent settings, as it enables independent agents to share information in partially observable settings and allows zero…

cs.AI20231 cited

Diverse Conventions for Human-AI Collaboration

Bidipta Sarkar, Andy Shih, Dorsa Sadigh

Conventions are crucial for strong performance in cooperative multi-agent games, because they allow players to coordinate on a shared strategy without explicit communication. Unfor…

cs.AI20237 cited

RoboCLIP: One Demonstration is Enough to Learn Robot Policies

Sumedh A Sontakke, Jesse Zhang, Sébastien M. R. Arnold +5

Reward specification is a notoriously difficult problem in reinforcement learning, requiring extensive expert supervision to design robust reward functions. Imitation learning (IL)…

cs.AI2023

Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Stephen Casper, Xander Davies, Claudia Shi +29

Reinforcement learning from human feedback (RLHF) is a technique for training AI systems to align with human goals. RLHF has emerged as the central method used to finetune state-of…