32 citations · 34 across the 4 of their papers we have counts for
4 papers
Reducing Human-Robot Goal State Divergence with Environment Design
Kelsey Sikes, Sarah Keren, Sarath Sreedharan
One of the most difficult challenges in creating successful human-AI collaborations is aligning a robot's behavior with a human user's expectations. When this fails to occur, a rob…
Planning for Attacker Entrapment in Adversarial Settings
Brittany Cates, Anagha Kulkarni, Sarath Sreedharan
In this paper, we propose a planning framework to generate a defense strategy against an attacker who is working in an environment where a defender can operate without the attacker…
On the Planning Abilities of Large Language Models (A Critical Investigation with a Proposed Benchmark)
Karthik Valmeekam, Sarath Sreedharan, Matthew Marquez +2
Intrigued by the claims of emergent reasoning capabilities in LLMs trained on general web corpora, in this paper, we set out to investigate their planning capabilities. We aim to e…
Goal Alignment: A Human-Aware Account of Value Alignment Problem
Malek Mechergui, Sarath Sreedharan
Value alignment problems arise in scenarios where the specified objectives of an AI agent don't match the true underlying objective of its users. The problem has been widely argued…