papers

Publications (14)

cs.CV2023

Navigating to Objects Specified by Images

Jacob Krantz, Theophile Gervet, Karmesh Yadav +7

Images are a convenient way to specify which particular object instance an embodied agent should navigate to. Solving this task requires semantic visual reasoning and exploration o…

cs.CV2021

Where Are You? Localization from Embodied Dialog

Meera Hahn, Jacob Krantz, Dhruv Batra +4

We present Where Are You? (WAY), a dataset of ~6k dialogs in which two humans -- an Observer and a Locator -- complete a cooperative localization task. The Observer is spawned at r…

cs.CV2022

Instance-Specific Image Goal Navigation: Training Embodied Agents to Find Object Instances

Jacob Krantz, Stefan Lee, Jitendra Malik +2

We consider the problem of embodied visual navigation given an image-goal (ImageNav) where an agent is initialized in an unfamiliar environment and tasked with navigating to a loca…

cs.CV2021

Waypoint Models for Instruction-guided Navigation in Continuous Environments

Jacob Krantz, Aaron Gokaslan, Dhruv Batra +2

Little inquiry has explicitly addressed the role of action spaces in language-guided visual navigation -- either in terms of its effect on navigation success or the efficiency with…

cs.CV2025

Do Visual Imaginations Improve Vision-and-Language Navigation Agents?

Akhil Perincherry, Jacob Krantz, Stefan Lee

Vision-and-Language Navigation (VLN) agents are tasked with navigating an unseen environment using natural language instructions. In this work, we study if visual representations o…

cs.RO2024

PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks

Matthew Chang, Gunjan Chhablani, Alexander Clegg +17

We present a benchmark for Planning And Reasoning Tasks in humaN-Robot collaboration (PARTNR) designed to study human-robot coordination in household activities. PARTNR tasks exhib…

cs.CV2023

Iterative Vision-and-Language Navigation

Jacob Krantz, Shurjo Banerjee, Wang Zhu +4

We present Iterative Vision-and-Language Navigation (IVLN), a paradigm for evaluating language-guided agents navigating in a persistent environment over time. Existing Vision-and-L…

cs.CV2022

Retrospectives on the Embodied AI Workshop

Matt Deitke, Dhruv Batra, Yonatan Bisk +36

We present a retrospective on the state of Embodied AI research. Our analysis focuses on 13 challenges presented at the Embodied AI Workshop at CVPR. These challenges are grouped i…

cs.CL2018

Abstractive Summarization Using Attentive Neural Techniques

Jacob Krantz, Jugal Kalita

In a world of proliferating data, the ability to rapidly summarize text is growing in importance. Automatic summarization of text can be thought of as a sequence to sequence proble…

cs.CL2019

Language-Agnostic Syllabification with Neural Sequence Labeling

Jacob Krantz, Maxwell Dulin, Paul De Palma

The identification of syllables within phonetic sequences is known as syllabification. This task is thought to play an important role in natural language understanding, speech prod…

cs.CV2022

Sim-2-Sim Transfer for Vision-and-Language Navigation in Continuous Environments

Jacob Krantz, Stefan Lee

Recent work in Vision-and-Language Navigation (VLN) has presented two environmental paradigms with differing realism -- the standard VLN setting built on topological environments w…

cs.CV2020

Beyond the Nav-Graph: Vision-and-Language Navigation in Continuous Environments

Jacob Krantz, Erik Wijmans, Arjun Majumdar +2

We develop a language-guided navigation task set in a continuous 3D environment where agents must execute low-level actions to follow natural language navigation directions. By bei…

math.DG2026

Partial Regularity of Stable Stationary Harmonic Maps into Certain Lie Groups

Jacob Krantz

Let be a compact Riemannian manifold, and let be a compact simple Lie group with bi-invariant metric that is not for , , , or…

cs.CL2018

Syllabification by Phone Categorization

Jacob Krantz, Maxwell Dulin, Paul De Palma +1

Syllables play an important role in speech synthesis, speech recognition, and spoken document retrieval. A novel, low cost, and language agnostic approach to dividing words into th…