activity
20142024
most citedDeep Captioning with Multimodal Recurrent Neural Networks (m-RNN)

652 citations · 1.7k across the 47 of their papers we have counts for

collaborators
Showing 2016Show all

5 papers · 1 filter

cs.LG201626 cited

Training and Evaluating Multimodal Word Embeddings with Large-scale Web Annotated Images

Junhua Mao, Jiajing Xu, Yushi Jing +1

In this paper, we focus on training and evaluating effective word embeddings with both text and visual information. More specifically, we introduce a large-scale dataset with 300 m…

cs.CV2016

Symmetric Non-Rigid Structure from Motion for Category-Specific Object Structure Estimation

Yuan Gao, Alan Yuille

Many objects, especially these made by humans, are symmetric, e.g. cars and aeroplanes. This paper addresses the estimation of 3D structures of symmetric objects from multiple imag…

cs.CV201614 cited

UnrealCV: Connecting Computer Vision to Unreal Engine

Weichao Qiu, Alan Yuille

Computer graphics can not only generate synthetic images and ground truth but it also offers the possibility of constructing virtual worlds in which: (i) an agent can perceive, nav…

cs.CV2016

Geometric Neural Phrase Pooling: Modeling the Spatial Co-occurrence of Neurons

Lingxi Xie, Qi Tian, John Flynn +2

Deep Convolutional Neural Networks (CNNs) are playing important roles in state-of-the-art visual recognition. This paper focuses on modeling the spatial co-occurrence of neuron res…

cs.CV20166 cited

Exploiting Symmetry and/or Manhattan Properties for 3D Object Structure Estimation from Single and Multiple Images

Yuan Gao, Alan L. Yuille

Many man-made objects have intrinsic symmetries and Manhattan structure. By assuming an orthographic projection model, this paper addresses the estimation of 3D structures and came…