Publications (14)
Energy-based Generative Adversarial Network
Junbo Zhao, Michael Mathieu, Yann LeCun
We introduce the "Energy-based Generative Adversarial Network" model (EBGAN) which views the discriminator as an energy function that attributes low energies to the regions near th…
Open-Ended Learning Leads to Generally Capable Agents
Open Ended Learning Team, Adam Stooke, Anuj Mahajan +15
In this work we create agents that can perform well beyond a single, individual task, that exhibit much wider generalisation of behaviour to a massive, rich space of challenges. We…
Disentangling factors of variation in deep representations using adversarial training
Michael Mathieu, Junbo Zhao, Pablo Sprechmann +2
We introduce a conditional generative model for learning to disentangle the hidden factors of variation within a set of labeled observations, and separate them into complementary c…
Learning to Linearize Under Uncertainty
Ross Goroshin, Michael Mathieu, Yann LeCun
Training deep feature hierarchies to solve supervised learning tasks has achieved state of the art performance on many problems in computer vision. However, a principled way in whi…
Learning Longer Memory in Recurrent Neural Networks
Tomas Mikolov, Armand Joulin, Sumit Chopra +2
Recurrent neural network is a powerful model that learns temporal patterns in sequential data. For a long time, it was believed that recurrent networks are difficult to train using…
Fast Convolutional Nets With fbfft: A GPU Performance Evaluation
Nicolas Vasilache, Jeff Johnson, Michael Mathieu +3
We examine the performance profile of Convolutional Neural Network training on the current generation of NVIDIA Graphics Processing Units. We introduce two new Fast Fourier Transfo…
Deep multi-scale video prediction beyond mean square error
Michael Mathieu, Camille Couprie, Yann LeCun
Learning to predict future images from a video sequence involves the construction of an internal representation that models the image evolution accurately, and therefore, to some d…
The Loss Surfaces of Multilayer Networks
Anna Choromanska, Mikael Henaff, Michael Mathieu +2
We study the connection between the highly non-convex loss function of a simple model of the fully-connected feed-forward neural network and the Hamiltonian of the spherical spin-g…
OverFeat: Integrated Recognition, Localization and Detection using Convolutional Networks
Pierre Sermanet, David Eigen, Xiang Zhang +3
We present an integrated framework for using Convolutional Networks for classification, localization and detection. We show how a multiscale and sliding window approach can be effi…
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431
In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our…
Stacked What-Where Auto-encoders
Junbo Zhao, Michael Mathieu, Ross Goroshin +1
We present a novel architecture, the "stacked what-where auto-encoders" (SWWAE), which integrates discriminative and generative pathways and provides a unified approach to supervis…
Fast Training of Convolutional Networks through FFTs
Michael Mathieu, Mikael Henaff, Yann LeCun
Convolutional networks are one of the most widely employed architectures in computer vision and machine learning. In order to leverage their ability to learn complex functions, lar…
Video (language) modeling: a baseline for generative models of natural videos
MarcAurelio Ranzato, Arthur Szlam, Joan Bruna +3
We propose a strong baseline model for unsupervised feature learning using video data. By learning to predict missing frames or extrapolate future frames from an input video sequen…
Fast Approximation of Rotations and Hessians matrices
Michael Mathieu, Yann LeCun
A new method to represent and approximate rotation matrices is introduced. The method represents approximations of a rotation matrix with linearithmic complexity, i.e. with $\f…