output
20052015
most citedBatch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift

24.4k citations

Showing 2014Show all

30 papers · 1 filter

cs.LG20143 cited

ACCAMS: Additive Co-Clustering to Approximate Matrices Succinctly

Alex Beutel, Amr Ahmed, Alexander J. Smola

Matrix completion and approximation are popular tools to capture a user's preferences for recommendation and to approximate missing data. Instead of using low-rank factorization we…

stat.ML201436 cited

The supervised hierarchical Dirichlet process

Andrew M. Dai, Amos J. Storkey

We propose the supervised hierarchical Dirichlet process (sHDP), a nonparametric generative model for the joint distribution of a group of observations and a response variable dire…

cs.LG2014699 cited

Multiple Object Recognition with Visual Attention

Jimmy Ba, Volodymyr Mnih, Koray Kavukcuoglu

We present an attention-based model for recognizing multiple objects in images. The proposed model is a deep recurrent neural network trained with reinforcement learning to attend…

cs.NE2014200 cited

Learning Longer Memory in Recurrent Neural Networks

Tomas Mikolov, Armand Joulin, Sumit Chopra +2

Recurrent neural network is a powerful model that learns temporal patterns in sequential data. For a long time, it was believed that recurrent networks are difficult to train using…

cs.NE201431 cited

Deep Networks With Large Output Spaces

Sudheendra Vijayanarasimhan, Jonathon Shlens, Rajat Monga +1

Deep neural networks have been extremely successful at various image, speech, video recognition tasks because of their ability to model deep structures within the data. However, th…

cs.CV2014119 cited

Attention for Fine-Grained Categorization

Pierre Sermanet, Andrea Frome, Esteban Real

This paper presents experiments extending the work of Ba et al. (2014) on recurrent neural models for attention into less constrained visual environments, specifically fine-grained…