DialFRED: Dialogue-Enabled Agents for Embodied Instruction Following
arXiv:2202.13330 · doi:10.1109/LRA.2022.3193254
Abstract
Language-guided Embodied AI benchmarks requiring an agent to navigate an environment and manipulate objects typically allow one-way communication: the human user gives a natural language command to the agent, and the agent can only follow the command passively. We present DialFRED, a dialogue-enabled embodied instruction following benchmark based on the ALFRED benchmark. DialFRED allows an agent to actively ask questions to the human user; the additional information in the user's response is used by the agent to better complete its task. We release a human-annotated dataset with 53K task-relevant questions and answers and an oracle to answer questions. To solve DialFRED, we propose a questioner-performer framework wherein the questioner is pre-trained with the human-annotated data and fine-tuned with reinforcement learning. We make DialFRED publicly available and encourage researchers to propose and evaluate their solutions to building dialog-enabled embodied agents.
8 pages, 5 figures, accepted by RA-L
References in corpus (4)
Cited by in corpus (6)
- Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions
- Can an Embodied Agent Find Your "Cat-shaped Mug"? LLM-Guided Exploration for Zero-Shot Object Navigation
- Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
- Learning to Generate Pointing Gestures in Situated Embodied Conversational Agents
- LEMMA: Learning Language-Conditioned Multi-Robot Manipulation
- HomeEmergency -- Using Audio to Find and Respond to Emergencies in the Home