Blindfold Baselines for Embodied QA
arXiv:1811.05013
Abstract
We explore blindfold (question-only) baselines for Embodied Question Answering. The EmbodiedQA task requires an agent to answer a question by intelligently navigating in a simulated environment, gathering necessary visual information only through first-person vision before finally answering. Consequently, a blindfold baseline which ignores the environment and visual information is a degenerate solution, yet we show through our experiments on the EQAv1 dataset that a simple question-only baseline achieves state-of-the-art results on the EmbodiedQA task in all cases except when the agent is spawned extremely close to the object.
NIPS 2018 Visually-Grounded Interaction and Language (ViGilL) Workshop
References in corpus (3)
Cited by in corpus (15)
- RUBi: Reducing Unimodal Biases in Visual Question Answering
- Don't Take the Easy Way Out: Ensemble Based Methods for Avoiding Known Dataset Biases
- Graph Density-Aware Losses for Novel Compositions in Scene Graph Generation
- Visual Dialogue without Vision or Dialogue
- Embodied Question Answering in Photorealistic Environments with Point Cloud Perception
- Deep Learning for Embodied Vision Navigation: A Survey
- Are You Looking? Grounding to Multiple Modalities in Vision-and-Language Navigation
- Interactive Language Learning by Question Answering
- VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering
- Multi-Target Embodied Question Answering
- Bayesian Relational Memory for Semantic Visual Navigation
- Misleading Failures of Partial-input Baselines
- Challenges and Prospects in Vision and Language Research
- A Revised Generative Evaluation of Visual Dialogue
- Why can't memory networks read effectively?