most citedMicrosoft COCO Captions: Data Collection and Evaluation Server

1.6k citations · 2.2k across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV201764 cited

Unifying Map and Landmark Based Representations for Visual Navigation

Saurabh Gupta, David Fouhey, Sergey Levine +1

This works presents a formulation for visual navigation that unifies map based spatial reasoning and path planning, with landmark based robust plan execution in noisy environments.…

cs.CV2015331 cited

Visual Semantic Role Labeling

Saurabh Gupta, Jitendra Malik

In this paper we introduce the problem of Visual Semantic Role Labeling: given an image we want to detect people doing actions and localize the objects of interaction. Classical ap…

cs.CV2015162 cited

Exploring Nearest Neighbor Approaches for Image Captioning

Jacob Devlin, Saurabh Gupta, Ross Girshick +2

We explore a variety of nearest neighbor baseline approaches for image captioning. These approaches find a set of nearest neighbor images in the training set from which a caption m…

cs.CV20151.6k cited

Microsoft COCO Captions: Data Collection and Evaluation Server

Xinlei Chen, Hao Fang, Tsung-Yi Lin +4

In this paper we describe the Microsoft COCO Caption dataset and evaluation server. When completed, the dataset will contain over one and a half million captions describing over 33…

cs.CV201530 cited

Inferring 3D Object Pose in RGB-D Images

Saurabh Gupta, Pablo Arbeláez, Ross Girshick +1

The goal of this work is to replace objects in an RGB-D scene with corresponding 3D models from a library. We approach this problem by first detecting and segmenting object instanc…