papers

Publications (54)

cs.AI2026

AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility

Xiaoyuan Liu, Jianhong Tu, Yuqi Chen +26

Agent systems are advancing quickly across domains, but their evaluation remains fragmented. Most benchmarks rely on fixed, LLM-centric harnesses that require heavy integration, cr…

cs.CR2025

DoomArena: A framework for Testing AI Agents Against Evolving Security Threats

Leo Boisvert, Mihir Bansal, Chandra Kiran Reddy Evuru +9

We present DoomArena, a security evaluation framework for AI agents. DoomArena is designed on three principles: 1) It is a plug-in framework and integrates easily into realistic ag…

stat.ML2017

PAC-Bayesian Theory Meets Bayesian Inference

Pascal Germain, Francis Bach, Alexandre Lacoste +1

We exhibit a strong link between frequentist PAC-Bayesian risk bounds and the Bayesian marginal likelihood. That is, for the negative log-likelihood loss function, we show that the…

cs.LG2021

Toward Foundation Models for Earth Monitoring: Proposal for a Climate Change Benchmark

Alexandre Lacoste, Evan David Sherwin, Hannah Kerner +9

Recent progress in self-supervision shows that pre-training large neural networks on vast amounts of unsupervised data can lead to impressive increases in generalisation for downst…

cs.LG2018

Improving Explorability in Variational Inference with Annealed Variational Objectives

Chin-Wei Huang, Shawn Tan, Alexandre Lacoste +1

Despite the advances in the representational capacity of approximate distributions for variational inference, the optimization process can still limit the density that is ultimatel…

stat.ML2024

Nonparametric Partial Disentanglement via Mechanism Sparsity: Sparse Actions, Interventions and Sparse Temporal Dependencies

Sébastien Lachapelle, Pau Rodríguez López, Yash Sharma +4

This work introduces a novel principle for disentanglement we call mechanism sparsity regularization, which applies when the latent factors of interest depend sparsely on observed…

cs.AI2025

InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation

Gaurav Sahu, Abhay Puri, Juan Rodriguez +11

Data analytics is essential for extracting valuable insights from data that can assist organizations in making effective decisions. We introduce InsightBench, a benchmark dataset w…

cs.LG2019

TADAM: Task dependent adaptive metric for improved few-shot learning

Boris N. Oreshkin, Pau Rodriguez, Alexandre Lacoste

Few-shot learning has become essential for producing models that generalize from few examples. In this work, we identify that metric scaling and metric task conditioning are import…

cs.CV2020

Embedding Propagation: Smoother Manifold for Few-Shot Classification

Pau Rodríguez, Issam Laradji, Alexandre Drouin +1

Few-shot classification is challenging because the data distribution of the training set can be widely different to the test set as their classes are disjoint. This distribution sh…

cs.LG2018

Neural Autoregressive Flows

Chin-Wei Huang, David Krueger, Alexandre Lacoste +1

Normalizing flows and autoregressive models have been successfully combined to produce state-of-the-art results in density estimation, via Masked Autoregressive Flows (MAF), and to…

cs.LG2025

Context is Key: A Benchmark for Forecasting with Essential Textual Information

Andrew Robert Williams, Arjun Ashok, Étienne Marcotte +8

Forecasting is a critical task in decision-making across numerous domains. While historical numerical data provide a start, they fail to convey the complete context for reliable an…

cs.AI2021

Online Fast Adaptation and Knowledge Accumulation: a New Approach to Continual Learning

Massimo Caccia, Pau Rodriguez, Oleksiy Ostapenko +8

Continual learning studies agents that learn from streams of tasks without forgetting previous ones while adapting to new ones. Two recent continual-learning scenarios have opened…

cs.LG2026

Privileged Information Distillation for Language Models

Emiliano Penaloza, Dheeraj Vattikonda, Nicolas Gontier +3

Training-time privileged information (PI) can enable language models to succeed on tasks they would otherwise fail, making it a powerful tool for reinforcement learning in hard, lo…

cs.LG2023

Capture the Flag: Uncovering Data Insights with Large Language Models

Issam Laradji, Perouz Taslakian, Sai Rajeswar +6

The extraction of a small number of relevant insights from vast amounts of data is a crucial component of data-driven decision-making. However, accomplishing this task requires con…

stat.ML2017

Deep Prior

Alexandre Lacoste, Thomas Boquet, Negar Rostamzadeh +3

The recent literature on deep learning offers new tools to learn a rich probability distribution over high dimensional data such as images or sounds. In this work we investigate th…

cs.LG2025

The BrowserGym Ecosystem for Web Agent Research

Thibault Le Sellier De Chezelles, Maxime Gasse, Alexandre Drouin +17

The BrowserGym ecosystem addresses the growing need for efficient evaluation and benchmarking of web agents, particularly those leveraging automation and Large Language Models (LLM…

cs.CY2019

Quantifying the Carbon Emissions of Machine Learning

Alexandre Lacoste, Alexandra Luccioni, Victor Schmidt +1

From an environmental standpoint, there are a few crucial aspects of training a neural network that have a major impact on the quantity of carbon that it emits. These factors inclu…

cs.CL2025

FocusAgent: Simple Yet Effective Ways of Trimming the Large Context of Web Agents

Imene Kerboua, Sahar Omidi Shayegan, Megh Thakkar +7

Web agents powered by large language models (LLMs) must process lengthy web page observations to complete user goals; these pages often exceed tens of thousands of tokens. This sat…

cs.LG2022

A General Purpose Neural Architecture for Geospatial Systems

Nasim Rahaman, Martin Weiss, Frederik Träuble +6

Geospatial Information Systems are used by researchers and Humanitarian Assistance and Disaster Response (HADR) practitioners to support a wide variety of important applications. H…

cs.CY2019

Tackling Climate Change with Machine Learning

David Rolnick, Priya L. Donti, Lynn H. Kaack +19

Climate change is one of the greatest challenges facing humanity, and we, as machine learning experts, may wonder how we can help. Here we describe how machine learning can be a po…

stat.ML2018

Bayesian Hypernetworks

David Krueger, Chin-Wei Huang, Riashat Islam +3

We study Bayesian hypernetworks: a framework for approximate Bayesian inference in neural networks. A Bayesian hypernetwork $\h$ is a neural network which learns to transform a sim…

cs.CR2026

Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain

Léo Boisvert, Léo Boisvert, Abhay Puri +8

While finetuning AI agents on interaction data -- such as web browsing or tool use -- improves their capabilities, it also introduces critical security vulnerabilities within the a…

cs.CV2020

Counting Cows: Tracking Illegal Cattle Ranching From High-Resolution Satellite Imagery

Issam Laradji, Pau Rodriguez, Freddie Kalaitzis +4

Cattle farming is responsible for 8.8\% of greenhouse gas emissions worldwide. In addition to the methane emitted due to their digestive process, the growing need for grazing areas…

cs.AI2025

WorkArena++: Towards Compositional Planning and Reasoning-based Common Knowledge Work Tasks

Léo Boisvert, Megh Thakkar, Maxime Gasse +6

The ability of large language models (LLMs) to mimic human-like intelligence has led to a surge in LLM-based autonomous agents. Though recent LLMs seem capable of planning and reas…

cs.LG2020

Bayesian active learning for production, a systematic study and a reusable library

Parmida Atighehchian, Frédéric Branchaud-Charron, Alexandre Lacoste

Active learning is able to reduce the amount of labelling effort by using a machine learning model to query the user for specific inputs. While there are many papers on new active…

cs.LG2020

Stochastic Neural Network with Kronecker Flow

Chin-Wei Huang, Ahmed Touati, Pascal Vincent +3

Recent advances in variational inference enable the modelling of highly structured joint distributions, but are limited in their capacity to scale to the high-dimensional setting o…

cs.LG2021

Beyond Trivial Counterfactual Explanations with Diverse Valuable Explanations

Pau Rodriguez, Massimo Caccia, Alexandre Lacoste +4

Explainability for machine learning models has gained considerable attention within the research community given the importance of deploying more reliable machine-learning systems.…

stat.ML2018

Uncertainty in Multitask Transfer Learning

Alexandre Lacoste, Boris Oreshkin, Wonchang Chung +3

Using variational Bayes neural networks, we develop an algorithm capable of accumulating knowledge into a prior from multiple different tasks. The result is a rich and meaningful p…

cs.LG2019

Hierarchical Importance Weighted Autoencoders

Chin-Wei Huang, Kris Sankaran, Eeshan Dhekane +2

Importance weighted variational inference (Burda et al., 2015) uses multiple i.i.d. samples to have a tighter variational lower bound. We believe a joint proposal has the potential…

cs.CL2017

WikiReading: A Novel Large-scale Language Understanding Task over Wikipedia

Daniel Hewlett, Alexandre Lacoste, Llion Jones +5

We present WikiReading, a large-scale natural language understanding task and publicly-available dataset with 18 million instances. The task is to predict textual values from the s…

cs.AI2024

Choreographer: Learning and Adapting Skills in Imagination

Pietro Mazzaglia, Tim Verbelen, Bart Dhoedt +2

Unsupervised skill learning aims to learn a rich repertoire of behaviors without external supervision, providing artificial agents with the ability to control and influence the env…

cs.CR2026

Indirect Prompt Injections: Are Firewalls All You Need, or Stronger Benchmarks?

Rishika Bhagwatkar, Kevin Kasa, Abhay Puri +5

AI agents are vulnerable to indirect prompt injection attacks, where malicious instructions embedded in external content or tool outputs cause unintended or harmful behavior. Inspi…

cs.AI2026

JEF-Hinter: Leveraging Offline Knowledge for Improving Web Agents Adaptation

Hadi Nekoei, Aman Jaiswal, Patrice Bechard +7

Large language model (LLM) agents perform well in sequential decision-making tasks, but improving them on unfamiliar domains often requires costly online interactions or fine-tunin…

stat.ML2022

Disentanglement via Mechanism Sparsity Regularization: A New Principle for Nonlinear ICA

Sébastien Lachapelle, Pau Rodríguez López, Yash Sharma +4

This work introduces a novel principle we call disentanglement via mechanism sparsity regularization, which can be applied when the latent factors of interest depend sparsely on pa…

cs.CL2026

Mem-: Adaptive Memory through Learning When and What to Generate

Xiaoqiang Wang, Chao Wang, Hadi Nekoei +5

We present Mem-, a framework for adaptive memory in large language model (LLM) agents, where useful guidance is generated on demand rather than retrieved from external memory s…

cs.LG2023

GEO-Bench: Toward Foundation Models for Earth Monitoring

Alexandre Lacoste, Nils Lehmann, Pau Rodriguez +14

Recent progress in self-supervision has shown that pre-training large neural networks on vast amounts of unsupervised data can lead to substantial increases in generalization to do…

cs.AI2026

How to Train Your LLM Web Agent: A Statistical Diagnosis

Dheeraj Vattikonda, Santhoshi Ravichandran, Emiliano Penaloza +13

LLM-based web agents have recently made significant progress, but much of it has occurred in closed-source systems, widening the gap with open-source alternatives. Progress has bee…

cs.AI2023

Mastering the Unsupervised Reinforcement Learning Benchmark from Pixels

Sai Rajeswar, Pietro Mazzaglia, Tim Verbelen +4

Controlling artificial agents from visual sensory data is an arduous task. Reinforcement learning (RL) algorithms can succeed but require large amounts of interactions between the…

cs.LG2022

Typing assumptions improve identification in causal discovery

Philippe Brouillard, Perouz Taslakian, Alexandre Lacoste +2

Causal discovery from observational data is a challenging task that can only be solved up to a set of equivalent solutions, called an equivalence class. Such classes, which are oft…

cs.CV2025

EarthView: A Large Scale Remote Sensing Dataset for Self-Supervision

Diego Velazquez, Pau Rodriguez López, Sergio Alonso +6

This paper presents EarthView, a comprehensive dataset specifically designed for self-supervision on remote sensing data, intended to enhance deep learning applications on Earth mo…

cs.LG2020

Differentiable Causal Discovery from Interventional Data

Philippe Brouillard, Sébastien Lachapelle, Alexandre Lacoste +2

Learning a causal directed acyclic graph from data is a challenging task that involves solving a combinatorial problem for which the solution is not always identifiable. A new line…

cs.AI2026

CUBE: A Standard for Unifying Agent Benchmarks

Alexandre Lacoste, Nicolas Gontier, Oleh Shliazhko +23

The proliferation of agent benchmarks has created critical fragmentation that threatens research productivity. Each new benchmark requires substantial custom integration, creating…

cs.CL2017

Hierarchical Question Answering for Long Documents

Eunsol Choi, Daniel Hewlett, Alexandre Lacoste +3

We present a framework for question answering that can efficiently scale to longer documents while maintaining or even improving performance of state-of-the-art models. While most…

cs.CL2026

Knowing When Not to Answer: Evaluating Abstention in Multimodal Reasoning Systems

Nishanth Madhusudhan, Vikas Yadav, Alexandre Lacoste

Effective abstention (EA), recognizing evidence insufficiency and refraining from answering, is critical for reliable multimodal systems. Yet existing evaluation paradigms for visi…

cs.LG2024

WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?

Alexandre Drouin, Maxime Gasse, Massimo Caccia +9

We study the use of large language model-based agents for interacting with software via web browsers. Unlike prior work, we focus on measuring the agents' ability to perform tasks…

cs.CV2025

GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks

Muhammad Sohail Danish, Muhammad Akhtar Munir, Syed Roshaan Ali Shah +5

While numerous recent benchmarks focus on evaluating generic Vision-Language Models (VLMs), they do not effectively address the specific challenges of geospatial applications. Gene…

cs.CV2021

Seasonal Contrast: Unsupervised Pre-Training from Uncurated Remote Sensing Data

Oscar Mañas, Alexandre Lacoste, Xavier Giro-i-Nieto +2

Remote sensing and automatic earth monitoring are key to solve global-scale challenges such as disaster prevention, land use monitoring, or tackling climate change. Although there…

cs.CL2025

LineRetriever: Planning-Aware Observation Reduction for Web Agents

Imene Kerboua, Sahar Omidi Shayegan, Megh Thakkar +6

While large language models have demonstrated impressive capabilities in web navigation tasks, the extensive context of web pages, often represented as DOM or Accessibility Tree (A…

cs.LG2021

Can Active Learning Preemptively Mitigate Fairness Issues?

Frédéric Branchaud-Charron, Parmida Atighehchian, Pau Rodríguez +2

Dataset bias is one of the prevailing causes of unfairness in machine learning. Addressing fairness at the data collection and dataset preparation stages therefore becomes an essen…

cs.LG2020

Adaptive Deep Kernel Learning

Prudencio Tossou, Basile Dura, Francois Laviolette +2

Deep kernel learning provides an elegant and principled framework for combining the structural properties of deep learning algorithms with the flexibility of kernel methods. By mea…

cs.LG2014

Sequential Model-Based Ensemble Optimization

Alexandre Lacoste, Hugo Larochelle, François Laviolette +1

One of the most tedious tasks in the application of machine learning is model selection, i.e. hyperparameter selection. Fortunately, recent progress has been made in the automation…

cs.CV2020

Synbols: Probing Learning Algorithms with Synthetic Datasets

Alexandre Lacoste, Pau Rodríguez, Frédéric Branchaud-Charron +7

Progress in the field of machine learning has been fueled by the introduction of benchmark datasets pushing the limits of existing algorithms. Enabling the design of datasets to te…

cs.CV2026

GEO-Bench-2: From Performance to Capability, Rethinking Evaluation in Geospatial AI

Naomi Simumba, Nils Lehmann, Paolo Fraccaro +9

Geospatial Foundation Models (GeoFMs) are transforming Earth Observation (EO), but evaluation lacks standardized protocols. GEO-Bench-2 addresses this with a comprehensive framewor…

cs.LG2021

Variational Causal Networks: Approximate Bayesian Inference over Causal Structures

Yashas Annadani, Jonas Rothfuss, Alexandre Lacoste +4

Learning the causal structure that underlies data is a crucial step towards robust real-world decision making. The majority of existing work in causal inference focuses on determin…