activity
20122024
most citedTarget-driven Visual Navigation in Indoor Scenes using Deep Reinforcement Learning

163 citations · 676 across the 18 of their papers we have counts for

collaborators
Showing 2016Show all

6 papers · 1 filter

cs.CV201646 cited

CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning

Justin Johnson, Bharath Hariharan, Laurens van der Maaten +3

When building artificial intelligence systems that can reason and answer questions about visual data, we need diagnostic tests to analyze our progress and discover shortcomings. Ex…

cs.CV20168 cited

Recurrent Attention Models for Depth-Based Person Identification

Albert Haque, Alexandre Alahi, Li Fei-Fei

We present an attention-based model that reasons on human body shape and motion dynamics to identify individuals in the absence of RGB information, hence in the dark. Our approach…

cs.HC201636 cited

A Glimpse Far into the Future: Understanding Long-term Crowd Worker Quality

Kenji Hata, Ranjay Krishna, Li Fei-Fei +1

Microtask crowdsourcing is increasingly critical to the creation of extremely large datasets. As a result, crowd workers spend weeks or months repeating the exact same tasks, makin…

cs.CV2016163 cited

Target-driven Visual Navigation in Indoor Scenes using Deep Reinforcement Learning

Yuke Zhu, Roozbeh Mottaghi, Eric Kolve +4

Two less addressed issues of deep reinforcement learning are (1) lack of generalization capability to new target goals, and (2) data inefficiency i.e., the model requires several (…

cs.CV2016111 cited

Visual Relationship Detection with Language Priors

Cewu Lu, Ranjay Krishna, Michael Bernstein +1

Visual relationships capture a wide variety of interactions between pairs of objects in images (e.g. "man riding bicycle" and "man pushing bicycle"). Consequently, the set of possi…

cs.CV201627 cited

Connectionist Temporal Modeling for Weakly Supervised Action Labeling

De-An Huang, Li Fei-Fei, Juan Carlos Niebles

We propose a weakly-supervised framework for action labeling in video, where only the order of occurring actions is required during training time. The key challenge is that the per…