papers

Publications (95)

stat.ML2020

Learning Neural Causal Models from Unknown Interventions

Nan Rosemary Ke, Olexa Bilaniuk, Anirudh Goyal +6

Promising results have driven a recent surge of interest in continuous optimization methods for Bayesian network structure learning from observational data. However, there are theo…

cs.AI2025

ReviewerToo: Should AI Join The Program Committee? A Look At The Future of Peer Review

Gaurav Sahu, Hugo Larochelle, Laurent Charlin +1

Peer review is the cornerstone of scientific publishing, yet it suffers from inconsistencies, reviewer subjectivity, and scalability challenges. We introduce ReviewerToo, a modular…

cs.LG2022

Head2Toe: Utilizing Intermediate Representations for Better Transfer Learning

Utku Evci, Vincent Dumoulin, Hugo Larochelle +1

Transfer-learning methods aim to improve performance in a data-scarce target domain using a model pretrained on a data-rich source domain. A cost-efficient strategy, linear probing…

cs.LG2022

Teaching Algorithmic Reasoning via In-context Learning

Hattie Zhou, Azade Nova, Hugo Larochelle +3

Large language models (LLMs) have shown increasing in-context learning capabilities through scaling up model and data size. Despite this progress, LLMs are still unable to solve al…

cs.LG2020

Learning to Execute Programs with Instruction Pointer Attention Graph Neural Networks

David Bieber, Charles Sutton, Hugo Larochelle +1

Graph neural networks (GNNs) have emerged as a powerful tool for learning software engineering tasks including code completion, bug finding, and program repair. They benefit from l…

cs.CL2014

Learning Multilingual Word Representations using a Bag-of-Words Autoencoder

Stanislas Lauly, Alex Boulanger, Hugo Larochelle

Recent work on learning multilingual word representations usually relies on the use of word-level alignements (e.g. infered with the help of GIZA++) between translated sentences, i…

cs.CV2025

Bringing SAM to new heights: Leveraging elevation data for tree crown segmentation from drone imagery

Mélisande Teng, Arthur Ouaknine, Etienne Laliberté +3

Information on trees at the individual level is crucial for monitoring forest ecosystems and planning forest management. Current monitoring methods involve ground measurements, req…

cs.LG2024

Many-Shot In-Context Learning

Rishabh Agarwal, Avi Singh, Lei M. Zhang +12

Large language models (LLMs) excel at few-shot in-context learning (ICL) -- learning from a few examples provided in context at inference, without any weight updates. Newly expande…

cs.CV2018

Blindfold Baselines for Embodied QA

Ankesh Anand, Eugene Belilovsky, Kyle Kastner +2

We explore blindfold (question-only) baselines for Embodied Question Answering. The EmbodiedQA task requires an agent to answer a question by intelligently navigating in a simulate…

cs.LG2021

Comparing Transfer and Meta Learning Approaches on a Unified Few-Shot Classification Benchmark

Vincent Dumoulin, Neil Houlsby, Utku Evci +4

Meta and transfer learning are two successful families of approaches to few-shot learning. Despite highly related goals, state-of-the-art advances in each family are measured large…

cs.LG2024

A density estimation perspective on learning from pairwise human preferences

Vincent Dumoulin, Daniel D. Johnson, Pablo Samuel Castro +2

Learning from human feedback (LHF) -- and in particular learning from pairwise preferences -- has recently become a crucial ingredient in training large language models (LLMs), and…

cs.LG2016

Document Neural Autoregressive Distribution Estimation

Stanislas Lauly, Yin Zheng, Alexandre Allauzen +1

We present an approach based on feed-forward neural networks for learning the distribution of textual documents. This approach is inspired by the Neural Autoregressive Distribution…

cs.LG2021

Learning a Universal Template for Few-shot Dataset Generalization

Eleni Triantafillou, Hugo Larochelle, Richard Zemel +1

Few-shot dataset generalization is a challenging variant of the well-studied few-shot classification problem where a diverse training set of several datasets is given, for the purp…

physics.optics2016

Deep Learning with Coherent Nanophotonic Circuits

Yichen Shen, Nicholas C. Harris, Scott Skirlo +8

Artificial Neural Networks are computational network models inspired by signal processing in the brain. These models have dramatically improved the performance of many learning tas…

cs.CV2015

Within-Brain Classification for Brain Tumor Segmentation

Mohammad Havaei, Hugo Larochelle, Philippe Poulin +1

Purpose: In this paper, we investigate a framework for interactive brain tumor segmentation which, at its core, treats the problem of interactive brain tumor segmentation as a mach…

cs.CV2017

Deep learning trends for focal brain pathology segmentation in MRI

Mohammad Havaei, Nicolas Guizard, Hugo Larochelle +1

Segmentation of focal (localized) brain pathologies such as brain tumors and brain lesions caused by multiple sclerosis and ischemic strokes are necessary for medical diagnosis, su…

cs.LG2022

Fortuitous Forgetting in Connectionist Networks

Hattie Zhou, Ankit Vani, Hugo Larochelle +1

Forgetting is often seen as an unwanted characteristic in both human and machine learning. However, we propose that forgetting can in fact be favorable to learning. We introduce "f…

cs.LG2025

CISO: Species Distribution Modeling Conditioned on Incomplete Species Observations

Hager Radi Abdelwahed, Mélisande Teng, Robin Zbinden +4

Species distribution models (SDMs) are widely used to predict species' geographic distributions, serving as critical tools for ecological research and conservation planning. Typica…

cs.LG2016

Dynamic Capacity Networks

Amjad Almahairi, Nicolas Ballas, Tim Cooijmans +3

We introduce the Dynamic Capacity Network (DCN), a neural network that can adaptively assign its capacity across different portions of the input data. This is achieved by combining…

cs.CL2015

Correlational Neural Networks

Sarath Chandar, Mitesh M. Khapra, Hugo Larochelle +1

Common Representation Learning (CRL), wherein different descriptions (or views) of the data are embedded in a common subspace, is receiving a lot of attention recently. Two popular…

cs.LG2016

An Infinite Restricted Boltzmann Machine

Marc-Alexandre Côté, Hugo Larochelle

We present a mathematical construction for the restricted Boltzmann machine (RBM) that doesn't require specifying the number of hidden units. In fact, the hidden layer size is adap…

cs.LG2021

Curriculum By Smoothing

Samarth Sinha, Animesh Garg, Hugo Larochelle

Convolutional Neural Networks (CNNs) have shown impressive performance in computer vision tasks such as image classification, detection, and segmentation. Moreover, recent work in…

cs.CV2016

Movie Description

Anna Rohrbach, Atousa Torabi, Marcus Rohrbach +5

Audio Description (AD) provides linguistic descriptions of movies and allows visually impaired people to follow a movie along with their peers. Such descriptions are by design main…

cs.LG2016

Autoencoding beyond pixels using a learned similarity metric

Anders Boesen Lindbo Larsen, Søren Kaae Sønderby, Hugo Larochelle +1

We present an autoencoder that leverages learned representations to better measure similarities in data space. By combining a variational autoencoder with a generative adversarial…

stat.ML2012

Practical Bayesian Optimization of Machine Learning Algorithms

Jasper Snoek, Hugo Larochelle, Ryan P. Adams

Machine learning algorithms frequently require careful tuning of model hyperparameters, regularization terms, and optimization parameters. Unfortunately, this tuning is often a "bl…

cs.LG2026

Towards Sustainable Investment Policies Informed by Opponent Shaping

Juan Agustin Duque, Razvan Ciuca, Ayoub Echchahed +2

Addressing climate change requires global coordination, yet rational economic actors often prioritize immediate gains over collective welfare, resulting in social dilemmas. InvestE…

cs.CV2013

A Supervised Neural Autoregressive Topic Model for Simultaneous Image Classification and Annotation

Yin Zheng, Yu-Jin Zhang, Hugo Larochelle

Topic modeling based on latent Dirichlet allocation (LDA) has been a framework of choice to perform scene recognition and annotation. Recently, a new type of topic model called the…

cs.CV2021

Self-Supervised Equivariant Scene Synthesis from Video

Cinjon Resnick, Or Litany, Cosmas Heiß +3

We propose a self-supervised framework to learn scene representations from video that are automatically delineated into background, characters, and their animations. Our method cap…

stat.ML2019

Small-GAN: Speeding Up GAN Training Using Core-sets

Samarth Sinha, Han Zhang, Anirudh Goyal +3

Recent work by Brock et al. (2018) suggests that Generative Adversarial Networks (GANs) benefit disproportionately from large mini-batch sizes. Unfortunately, using large batches i…

cs.CV2023

Bird Distribution Modelling using Remote Sensing and Citizen Science data

Mélisande Teng, Amna Elmustafa, Benjamin Akera +2

Climate change is a major driver of biodiversity loss, changing the geographic range and abundance of many species. However, there remain significant knowledge gaps about the distr…

stat.ML2015

Describing Videos by Exploiting Temporal Structure

Li Yao, Atousa Torabi, Kyunghyun Cho +4

Recent progress in using recurrent neural networks (RNNs) for image description has motivated the exploration of their application for video description. However, while images are…

cs.CV2021

Impact of Aliasing on Generalization in Deep Convolutional Networks

Cristina Vasconcelos, Hugo Larochelle, Vincent Dumoulin +3

We investigate the impact of aliasing on generalization in Deep Convolutional Networks and show that data augmentation schemes alone are unable to prevent it due to structural limi…

cs.SD2025

Identifying birdsong syllables without labelled data

Mélisande Teng, Julien Boussard, David Rolnick +1

Identifying sequences of syllables within birdsongs is key to tackling a wide array of challenges, including bird individual identification and better understanding of animal commu…

cs.LG2019

The Hanabi Challenge: A New Frontier for AI Research

Nolan Bard, Jakob N. Foerster, Sarath Chandar +12

From the early days of computing, games have been important testbeds for studying how well machines can do sophisticated decision making. In recent years, machine learning has made…

stat.ML2023

InfoBot: Transfer and Exploration via the Information Bottleneck

Anirudh Goyal, Riashat Islam, Daniel Strouse +5

A central challenge in reinforcement learning is discovering effective policies for tasks where rewards are sparsely distributed. We postulate that in the absence of useful reward…

cs.LG2021

Your GAN is Secretly an Energy-based Model and You Should use Discriminator Driven Latent Sampling

Tong Che, Ruixiang Zhang, Jascha Sohl-Dickstein +4

We show that the sum of the implicit generator log-density of a GAN with the logit score of the discriminator defines an energy function which yields the true data densi…

stat.ML2016

Domain-Adversarial Training of Neural Networks

Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan +5

We introduce a new representation learning approach for domain adaptation, in which data at training and test time come from similar but different distributions. Our approach is di…

cs.CV2015

A Deep and Autoregressive Approach for Topic Modeling of Multimodal Data

Yin Zheng, Yu-Jin Zhang, Hugo Larochelle

Topic modeling based on latent Dirichlet allocation (LDA) has been a framework of choice to deal with multimodal data, such as in image annotation tasks. Another popular approach t…

cs.CV2025

Assessing SAM for Tree Crown Instance Segmentation from Drone Imagery

Mélisande Teng, Arthur Ouaknine, Etienne Laliberté +3

The potential of tree planting as a natural climate solution is often undermined by inadequate monitoring of tree planting projects. Current monitoring methods involve measuring tr…

cs.LG2020

Revisiting Fundamentals of Experience Replay

William Fedus, Prajit Ramachandran, Rishabh Agarwal +4

Experience replay is central to off-policy algorithms in deep reinforcement learning (RL), but there remain significant gaps in our understanding. We therefore present a systematic…

cs.LG2020

Diversity inducing Information Bottleneck in Model Ensembles

Samarth Sinha, Homanga Bharadhwaj, Anirudh Goyal +3

Although deep learning models have achieved state-of-the-art performance on a number of vision tasks, generalization over high dimensional multi-modal data, and reliable predictive…

cs.CV2020

An Effective Anti-Aliasing Approach for Residual Networks

Cristina Vasconcelos, Hugo Larochelle, Vincent Dumoulin +2

Image pre-processing in the frequency domain has traditionally played a vital role in computer vision and was even part of the standard pipeline in the early days of deep learning.…

cs.LG2020

Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few Examples

Eleni Triantafillou, Tyler Zhu, Vincent Dumoulin +8

Few-shot classification refers to learning a classifier for new classes given only a few examples. While a plethora of models have emerged to tackle it, we find the procedure and d…

cs.CV2016

Brain Tumor Segmentation with Deep Neural Networks

Mohammad Havaei, Axel Davy, David Warde-Farley +6

In this paper, we present a fully automatic brain tumor segmentation method based on Deep Neural Networks (DNNs). The proposed networks are tailored to glioblastomas (both low and…

stat.ML2016

Hierarchical Memory Networks

Sarath Chandar, Sungjin Ahn, Hugo Larochelle +3

Memory networks are neural networks with an explicit memory component that can be both read and written to by the network. The memory is often addressed in a soft way using a softm…

stat.ML2014

A Deep and Tractable Density Estimator

Benigno Uria, Iain Murray, Hugo Larochelle

The Neural Autoregressive Distribution Estimator (NADE) and its real-valued version RNADE are competitive density models of multidimensional data across a variety of domains. These…

cs.LG2023

Repository-Level Prompt Generation for Large Language Models of Code

Disha Shrivastava, Hugo Larochelle, Daniel Tarlow

With the success of large language models (LLMs) of code and their use as code assistants (e.g. Codex used in GitHub Copilot), techniques for introducing domain-specific knowledge…

cs.LG2016

Neural Autoregressive Distribution Estimation

Benigno Uria, Marc-Alexandre Côté, Karol Gregor +2

We present Neural Autoregressive Distribution Estimation (NADE) models, which are neural network architectures applied to the problem of unsupervised distribution and density estim…

cs.AI2026

The Alien Space of Science: Sampling Coherent but Cognitively Unavailable Research Directions

Alejandro H. Artiles, Martin Weiss, Levin Brinkmann +6

Scientific discovery is constrained not only by what is true, but by what is cognitively available to the researchers currently exploring a field. Many directions are coherent in l…

stat.ML2011

Loss-sensitive Training of Probabilistic Conditional Random Fields

Maksims N. Volkovs, Hugo Larochelle, Richard S. Zemel

We consider the problem of training probabilistic conditional random fields (CRFs) in the context of a task where performance is measured using a specific loss function. While maxi…

cs.LG2026

Detoxifying LLMs via Representation Erasure-Based Preference Optimization

Nazanin Mohammadi Sepahvand, Eleni Triantafillou, Hugo Larochelle +3

Large language models (LLMs) trained on webscale data can produce toxic outputs, raising concerns for safe deployment. Prior defenses, based on applications of DPO, NPO, and simila…

cs.LG2020

Improving Reproducibility in Machine Learning Research (A Report from the NeurIPS 2019 Reproducibility Program)

Joelle Pineau, Philippe Vincent-Lamarre, Koustuv Sinha +5

One of the challenges in machine learning research is to ensure that presented and published results are sound and reliable. Reproducibility, that is obtaining similar results as p…

cs.LG2025

Don't flatten, tokenize! Unlocking the key to SoftMoE's efficacy in deep RL

Ghada Sokar, Johan Obando-Ceron, Aaron Courville +2

The use of deep neural networks in reinforcement learning (RL) often suffers from performance degradation as model size increases. While soft mixtures of experts (SoftMoEs) have re…

cs.AI2019

Algorithmic Improvements for Deep Reinforcement Learning applied to Interactive Fiction

Vishal Jain, William Fedus, Hugo Larochelle +2

Text-based games are a natural challenge domain for deep reinforcement learning algorithms. Their state and action spaces are combinatorially large, their reward function is sparse…

cs.AI2026

Capturing Individual Human Preferences with Reward Features

André Barreto, Vincent Dumoulin, Yiran Mao +6

Reinforcement learning from human feedback usually models preferences using a reward function that does not distinguish between people. We argue that this is unlikely to be a good…

cs.LG2019

Recall Traces: Backtracking Models for Efficient Reinforcement Learning

Anirudh Goyal, Philemon Brakel, William Fedus +5

In many environments only a tiny subset of all states yield high reward. In these cases, few of the interactions with the environment provide a relevant learning signal. Hence, we…

cs.LG2012

Conditional Restricted Boltzmann Machines for Structured Output Prediction

Volodymyr Mnih, Hugo Larochelle, Geoffrey E. Hinton

Conditional Restricted Boltzmann Machines (CRBMs) are rich probabilistic models that have recently been applied to a wide range of problems, including collaborative filtering, clas…

stat.ML2015

Domain-Adversarial Neural Networks

Hana Ajakan, Pascal Germain, Hugo Larochelle +2

We introduce a new representation learning algorithm suited to the context of domain adaptation, in which data at training and test time come from similar but different distributio…

eess.AS2025

The Search for Squawk: Agile Modeling in Bioacoustics

Vincent Dumoulin, Otilia Stretcu, Jenny Hamer +16

Passive acoustic monitoring (PAM) has shown great promise in helping ecologists understand the health of animal populations and ecosystems. However, extracting insights from millio…

cs.CV2022

Matching Feature Sets for Few-Shot Image Classification

Arman Afrasiyabi, Hugo Larochelle, Jean-François Lalonde +1

In image classification, it is common practice to train deep networks to extract a single feature vector per input image. Few-shot classification methods also mostly follow this tr…

cs.LG2011

Classification of Sets using Restricted Boltzmann Machines

Jérôme Louradour, Hugo Larochelle

We consider the problem of classification when inputs correspond to sets of vectors. This setting occurs in many problems such as the classification of pieces of mail containing se…

cs.LG2020

On-the-Fly Adaptation of Source Code Models using Meta-Learning

Disha Shrivastava, Hugo Larochelle, Daniel Tarlow

The ability to adapt to unseen, local contexts is an important challenge that successful models of source code must overcome. One of the most popular approaches for the adaptation…

cs.AI2011

Learning where to Attend with Deep Architectures for Image Tracking

Misha Denil, Loris Bazzani, Hugo Larochelle +1

We discuss an attentional model for simultaneous object tracking and recognition that is driven by gaze data. Motivated by theories of perception, the model consists of two interac…

cs.LG2022

Static Prediction of Runtime Errors by Learning to Execute Programs with External Resource Descriptions

David Bieber, Rishab Goel, Daniel Zheng +2

The execution behavior of a program often depends on external resources, such as program inputs or file contents, and so cannot be run in isolation. Nevertheless, software develope…

cs.LG2020

On Catastrophic Interference in Atari 2600 Games

William Fedus, Dibya Ghosh, John D. Martin +3

Model-free deep reinforcement learning is sample inefficient. One hypothesis -- speculated, but not confirmed -- is that catastrophic interference within an environment inhibits le…

cs.LG2018

Meta-Learning for Semi-Supervised Few-Shot Classification

Mengye Ren, Eleni Triantafillou, Sachin Ravi +5

In few-shot classification, we are interested in learning algorithms that train a classifier from only a handful of labeled examples. Recent progress in few-shot classification has…

stat.ML2011

On Nonparametric Guidance for Learning Autoencoder Representations

Jasper Snoek, Ryan Prescott Adams, Hugo Larochelle

Unsupervised discovery of latent representations, in addition to being useful for density modeling, visualisation and exploratory data analysis, is also increasingly important for…

cs.LG2020

Uniform Priors for Data-Efficient Transfer

Samarth Sinha, Karsten Roth, Anirudh Goyal +3

Deep Neural Networks have shown great promise on a variety of downstream applications; but their ability to adapt and generalize to new data and tasks remains a challenge. However,…

cs.LG2014

Sequential Model-Based Ensemble Optimization

Alexandre Lacoste, Hugo Larochelle, François Laviolette +1

One of the most tedious tasks in the application of machine learning is model selection, i.e. hyperparameter selection. Fortunately, recent progress has been made in the automation…

cs.LG2012

Training Restricted Boltzmann Machines on Word Observations

George E. Dahl, Ryan P. Adams, Hugo Larochelle

The restricted Boltzmann machine (RBM) is a flexible tool for modeling complex data, however there have been significant computational difficulties in using RBMs to model high-dime…

cs.LG2021

Learning to Combine Per-Example Solutions for Neural Program Synthesis

Disha Shrivastava, Hugo Larochelle, Daniel Tarlow

The goal of program synthesis from examples is to find a computer program that is consistent with a given set of input-output examples. Most learning-based approaches try to find a…

cs.LG2020

Are Few-Shot Learning Benchmarks too Simple ? Solving them without Task Supervision at Test-Time

Gabriel Huang, Hugo Larochelle, Simon Lacoste-Julien

We show that several popular few-shot learning benchmarks can be solved with varying degrees of success without using support set Labels at Test-time (LT). To this end, we introduc…

stat.ML2019

Hyperbolic Discounting and Learning over Multiple Horizons

William Fedus, Carles Gelada, Yoshua Bengio +2

Reinforcement learning (RL) typically defines a discount factor as part of the Markov Decision Process. The discount factor values future rewards by an exponential scheme that lead…

cs.AI2017

GuessWhat?! Visual object discovery through multi-modal dialogue

Harm de Vries, Florian Strub, Sarath Chandar +3

We introduce GuessWhat?!, a two-player guessing game as a testbed for research on the interplay of computer vision and dialogue systems. The goal of the game is to locate an unknow…

cs.LG2020

A RAD approach to deep mixture models

Laurent Dinh, Jascha Sohl-Dickstein, Hugo Larochelle +1

Flow based models such as Real NVP are an extremely powerful approach to density estimation. However, existing flow based models are restricted to transforming continuous densities…

cs.CV2017

Recurrent Mixture Density Network for Spatiotemporal Visual Attention

Loris Bazzani, Hugo Larochelle, Lorenzo Torresani

In many computer vision tasks, the relevant information to solve the problem at hand is mixed to irrelevant, distracting information. This has motivated researchers to design atten…

cs.LG2015

Clustering is Efficient for Approximate Maximum Inner Product Search

Alex Auvolat, Sarath Chandar, Pascal Vincent +2

Efficient Maximum Inner Product Search (MIPS) is an important task that has a wide applicability in recommendation systems and classification with a large number of classes. Soluti…

cs.LG2023

SatBird: Bird Species Distribution Modeling with Remote Sensing and Citizen Science Data

Mélisande Teng, Amna Elmustafa, Benjamin Akera +4

Biodiversity is declining at an unprecedented rate, impacting ecosystem services necessary to ensure food, water, and human health and well-being. Understanding the distribution of…

cs.CV2017

Modulating early visual processing by language

Harm de Vries, Florian Strub, Jérémie Mary +3

It is commonly assumed that language refers to high-level visual concepts while leaving low-level visual processing unaffected. This view dominates the current literature in comput…

cs.CL2014

An Autoencoder Approach to Learning Bilingual Word Representations

Sarath Chandar A P, Stanislas Lauly, Hugo Larochelle +4

Cross-language learning allows us to use training data from one language to build models for a different language. Many approaches to bilingual learning require that we have word-l…

cs.LG2011

Autotagging music with conditional restricted Boltzmann machines

Michael Mandel, Razvan Pascanu, Hugo Larochelle +1

This paper describes two applications of conditional restricted Boltzmann machines (CRBMs) to the task of autotagging music. The first consists of training a CRBM to predict tags t…

cs.AI2026

BRIDGE: Predicting Human Task Completion Time From Model Performance

Fengyuan Liu, Jay Gala, Nilaksh +3

Evaluating the real-world capabilities of AI systems requires grounding benchmark performance in human-interpretable measures of task difficulty. Existing approaches that rely on d…

cs.DL2026

Lacuna: A Research Map for Machine Learning

Martin Weiss, Miles Q. Li, Alejandro H. Artiles +4

Lacuna is a research map for machine learning that uses LLMs to turn papers and scholarly metadata into markdown summaries, concept elements, research directions, and research prop…

cs.LG2020

A Universal Representation Transformer Layer for Few-Shot Image Classification

Lu Liu, William Hamilton, Guodong Long +2

Few-shot classification aims to recognize unseen classes when presented with only a small number of samples. We consider the problem of multi-domain few-shot image classification,…

cs.CL2020

Language GANs Falling Short

Massimo Caccia, Lucas Caccia, William Fedus +3

Generating high-quality text with sufficient diversity is essential for a wide range of Natural Language Generation (NLG) tasks. Maximum-Likelihood (MLE) models trained with teache…

stat.ML2018

Disentangling the independently controllable factors of variation by interacting with the world

Valentin Thomas, Emmanuel Bengio, William Fedus +6

It has been postulated that a good representation is one that disentangles the underlying explanatory factors of variation. However, it remains an open question what kind of traini…

stat.ML2014

RNADE: The real-valued neural autoregressive density-estimator

Benigno Uria, Iain Murray, Hugo Larochelle

We introduce RNADE, a new model for joint density estimation of real-valued vectors. Our model calculates the density of a datapoint as the product of one-dimensional conditionals…

cs.LG2020

Learning Graph Structure With A Finite-State Automaton Layer

Daniel D. Johnson, Hugo Larochelle, Daniel Tarlow

Graph-based neural network models are producing strong results in a number of domains, in part because graphs provide flexibility to encode domain knowledge in the form of relation…

cs.CV2015

Using Descriptive Video Services to Create a Large Data Source for Video Annotation Research

Atousa Torabi, Christopher Pal, Hugo Larochelle +1

In this work, we introduce a dataset of video annotated with high quality natural language phrases describing the visual content in a given segment of time. Our dataset is based on…

cs.LG2024

Unlearning via Sparse Representations

Vedant Shah, Frederik Träuble, Ashish Malik +5

Machine \emph{unlearning}, which involves erasing knowledge about a \emph{forget set} from a trained model, can prove to be costly and infeasible by existing techniques. We propose…

cs.CL2021

Interpretable Multi-Modal Hate Speech Detection

Prashanth Vijayaraghavan, Hugo Larochelle, Deb Roy

With growing role of social media in shaping public opinions and beliefs across the world, there has been an increased attention to identify and counter the problem of hate speech…

cs.AI2017

HoME: a Household Multimodal Environment

Simon Brodeur, Ethan Perez, Ankesh Anand +6

We introduce HoME: a Household Multimodal Environment for artificial agents to learn from vision, audio, semantics, physics, and interaction with objects and other agents, all with…

stat.ML2017

Multiscale sequence modeling with a learned dictionary

Bart van Merriënboer, Amartya Sanyal, Hugo Larochelle +1

We propose a generalization of neural network sequence models. Instead of predicting one symbol at a time, our multi-scale model makes predictions over multiple, potentially overla…

cs.CV2020

Learned Equivariant Rendering without Transformation Supervision

Cinjon Resnick, Or Litany, Hugo Larochelle +2

We propose a self-supervised framework to learn scene representations from video that are automatically delineated into objects and background. Our method relies on moving objects…

cs.LG2015

MADE: Masked Autoencoder for Distribution Estimation

Mathieu Germain, Karol Gregor, Iain Murray +1

There has been a lot of recent interest in designing neural network models to estimate a distribution from a set of examples. We introduce a simple modification for autoencoder neu…