107 citations · 407 across the 30 of their papers we have counts for
4 papers · 2 filters
Red Teaming with Mind Reading: White-Box Adversarial Policies Against RL Agents
Stephen Casper, Taylor Killian, Gabriel Kreiman +1
Adversarial examples can be useful for identifying vulnerabilities in AI systems before they are deployed. In reinforcement learning (RL), adversarial policies can be developed by…
Formal Contracts Mitigate Social Dilemmas in Multi-Agent RL
Andreas A. Haupt, Phillip J. K. Christoffersen, Mehul Damani +1
Multi-agent Reinforcement Learning (MARL) is a powerful tool for training autonomous agents acting independently in a common environment. However, it can lead to sub-optimal behavi…
How to talk so AI will learn: Instructions, descriptions, and autonomy
Theodore R Sumers, Robert D Hawkins, Mark K Ho +2
From the earliest years of our lives, humans use language to express our beliefs and desires. Being able to talk to artificial agents about our preferences would thus fulfill a cen…
Linguistic communication as (inverse) reward design
Theodore R. Sumers, Robert D. Hawkins, Mark K. Ho +2
Natural language is an intuitive and expressive way to communicate reward information to autonomous agents. It encompasses everything from concrete instructions to abstract descrip…