24 citations · 46 across the 8 of their papers we have counts for
7 papers · 1 filter
FindingDory: A Benchmark to Evaluate Memory in Embodied Agents
Karmesh Yadav, Yusuf Ali, Gunshi Gupta +2
Large vision-language models have recently demonstrated impressive performance in planning and control tasks, driving interest in their application to real-world robotics. However,…
Pre-trained Text-to-Image Diffusion Models Are Versatile Representation Learners for Control
Gunshi Gupta, Karmesh Yadav, Yarin Gal +4
Embodied AI agents require a fine-grained understanding of the physical world mediated through visual and language inputs. Such capabilities are difficult to learn solely from task…
Navigating to Objects Specified by Images
Jacob Krantz, Theophile Gervet, Karmesh Yadav +7
Images are a convenient way to specify which particular object instance an embodied agent should navigate to. Solving this task requires semantic visual reasoning and exploration o…
OVRL-V2: A simple state-of-art baseline for ImageNav and ObjectNav
Karmesh Yadav, Arjun Majumdar, Ram Ramrakhya +5
We present a single neural network architecture composed of task-agnostic components (ViTs, convolutions, and LSTMs) that achieves state-of-art results on both the ImageNav ("go to…
Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?
Arjun Majumdar, Karmesh Yadav, Sergio Arnaud +12
We present the largest and most comprehensive empirical study of pre-trained visual representations (PVRs) or visual 'foundation models' for Embodied AI. First, we curate CortexBen…
Last-Mile Embodied Visual Navigation
Justin Wasserman, Karmesh Yadav, Girish Chowdhary +2
Realistic long-horizon tasks like image-goal navigation involve exploratory and exploitative phases. Assigned with an image of the goal, an embodied agent must explore to discover…