Publications (46)
Learning interpretable models of phenotypes from whole genome sequences with the Set Covering Machine
Alexandre Drouin, Sébastien Giguère, Vladana Sagatovich +4
The increased affordability of whole genome sequencing has motivated its use for phenotypic studies. We address the problem of learning interpretable models for discrete phenotypes…
TACTiS-2: Better, Faster, Simpler Attentional Copulas for Multivariate Time Series
Arjun Ashok, Ãtienne Marcotte, Valentina Zantedeschi +2
We introduce a new model for multivariate probabilistic time series prediction, designed to flexibly address a range of tasks including forecasting, interpolation, and their combin…
Maximum Margin Interval Trees
Alexandre Drouin, Toby Dylan Hocking, François Laviolette
Learning a regression function using censored or interval-valued output data is an important problem in fields such as genomics and medicine. The goal is to learn a real-valued pre…
Generalization Bounds via Meta-Learned Model Representations: PAC-Bayes and Sample Compression Hypernetworks
Benjamin Leblanc, Mathieu Bazinet, Nathaniel D'Amours +2
Both PAC-Bayesian and Sample Compress learning frameworks are instrumental for deriving tight (non-vacuous) generalization bounds for neural networks. We leverage these results in…
JEF-Hinter: Leveraging Offline Knowledge for Improving Web Agents Adaptation
Hadi Nekoei, Aman Jaiswal, Patrice Bechard +7
Large language model (LLM) agents perform well in sequential decision-making tasks, but improving them on unfamiliar domains often requires costly online interactions or fine-tunin…
Regions of Reliability in the Evaluation of Multivariate Probabilistic Forecasts
Ãtienne Marcotte, Valentina Zantedeschi, Alexandre Drouin +1
Multivariate probabilistic time series forecasts are commonly evaluated via proper scoring rules, i.e., functions that are minimal in expectation for the ground-truth distribution.…
GEO-Bench: Toward Foundation Models for Earth Monitoring
Alexandre Lacoste, Nils Lehmann, Pau Rodriguez +14
Recent progress in self-supervision has shown that pre-training large neural networks on vast amounts of unsupervised data can lead to substantial increases in generalization to do…
How to Train Your LLM Web Agent: A Statistical Diagnosis
Dheeraj Vattikonda, Santhoshi Ravichandran, Emiliano Penaloza +13
LLM-based web agents have recently made significant progress, but much of it has occurred in closed-source systems, widening the gap with open-source alternatives. Progress has bee…
Learning to Defer for Causal Discovery with Imperfect Experts
Oscar Clivio, Divyat Mahajan, Perouz Taslakian +4
Integrating expert knowledge, e.g. from large language models, into causal discovery algorithms can be challenging when the knowledge is not guaranteed to be correct. Expert recomm…
Large scale modeling of antimicrobial resistance with interpretable classifiers
Alexandre Drouin, Frédéric Raymond, Gaël Letarte St-Pierre +3
Antimicrobial resistance is an important public health concern that has implications in the practice of medicine worldwide. Accurately predicting resistance phenotypes from genome…
Typing assumptions improve identification in causal discovery
Philippe Brouillard, Perouz Taslakian, Alexandre Lacoste +2
Causal discovery from observational data is a challenging task that can only be solved up to a set of equivalent solutions, called an equivalence class. Such classes, which are oft…
The Landscape of Causal Discovery Data: Grounding Causal Discovery in Real-World Applications
Philippe Brouillard, Chandler Squires, Jonas Wahl +4
Causal discovery aims to automatically uncover causal relationships from data, a capability with significant potential across many scientific disciplines. However, its real-world a…
Differentiable Causal Discovery from Interventional Data
Philippe Brouillard, Sébastien Lachapelle, Alexandre Lacoste +2
Learning a causal directed acyclic graph from data is a challenging task that involves solving a combinatorial problem for which the solution is not always identifiable. A new line…
CUBE: A Standard for Unifying Agent Benchmarks
Alexandre Lacoste, Nicolas Gontier, Oleh Shliazhko +23
The proliferation of agent benchmarks has created critical fragmentation that threatens research productivity. Each new benchmark requires substantial custom integration, creating…
Beyond Naïve Prompting: Strategies for Improved Context-aided Forecasting with LLMs
Arjun Ashok, Andrew Robert Williams, Vincent Zhihao Zheng +5
Real-world forecasting requires models to integrate not only historical data but also relevant contextual information provided in textual form. While large language models (LLMs) s…
Overcoming the Modality Gap in Context-Aided Forecasting
Vincent Zhihao Zheng, Ãtienne Marcotte, Arjun Ashok +4
The paper introduces a semi‑synthetic data augmentation technique to create high‑quality contextual information for time‑series forecasting, producing a 7 million‑sample dataset (C…
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
Alexandre Drouin, Maxime Gasse, Massimo Caccia +9
We study the use of large language model-based agents for interacting with software via web browsers. Unlike prior work, we focus on measuring the agents' ability to perform tasks…
Apriel-1.5-OpenReasoner: RL Post-Training for General-Purpose and Efficient Reasoning
Rafael Pardinas, Ehsan Kamalloo, David Vazquez +1
Building general-purpose reasoning models using reinforcement learning with verifiable rewards (RLVR) across diverse domains has been widely adopted by frontier open-weight models.…
Learning a peptide-protein binding affinity predictor with kernel ridge regression
Sébastien Giguère, Mario Marchand, François Laviolette +2
We propose a specialized string kernel for small bio-molecules, peptides and pseudo-sequences of binding interfaces. The kernel incorporates physico-chemical properties of amino ac…
TACTiS: Transformer-Attentional Copulas for Time Series
Alexandre Drouin, Ãtienne Marcotte, Nicolas Chapados
The estimation of time-varying quantities is a fundamental component of decision making in fields such as healthcare and finance. However, the practical utility of such estimates i…
Evaluating Interventional Reasoning Capabilities of Large Language Models
Tejas Kasetty, Divyat Mahajan, Gintare Karolina Dziugaite +2
Numerous decision-making tasks require estimating causal effects under interventions on different parts of a system. As practitioners consider using large language models (LLMs) to…
Lag-Llama: Towards Foundation Models for Probabilistic Time Series Forecasting
Kashif Rasul, Arjun Ashok, Andrew Robert Williams +15
Over the past years, foundation models have caused a paradigm shift in machine learning due to their unprecedented capabilities for zero-shot and few-shot generalization. However,…
Causal Discovery with Language Models as Imperfect Experts
Stephanie Long, Alexandre Piché, Valentina Zantedeschi +2
Understanding the causal relationships that underlie a system is a fundamental prerequisite to accurate decision-making. In this work, we explore how expert knowledge can be used t…
Greedy Biomarker Discovery in the Genome with Applications to Antimicrobial Resistance
Alexandre Drouin, Sébastien Giguère, Maxime Déraspe +3
The Set Covering Machine (SCM) is a greedy learning algorithm that produces sparse classifiers. We extend the SCM for datasets that contain a huge number of features. The whole gen…
Synbols: Probing Learning Algorithms with Synthetic Datasets
Alexandre Lacoste, Pau RodrÃguez, Frédéric Branchaud-Charron +7
Progress in the field of machine learning has been fueled by the introduction of benchmark datasets pushing the limits of existing algorithms. Enabling the design of datasets to te…
AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility
Xiaoyuan Liu, Jianhong Tu, Yuqi Chen +26
Agent systems are advancing quickly across domains, but their evaluation remains fragmented. Most benchmarks rely on fixed, LLM-centric harnesses that require heavy integration, cr…
byteSteady: Fast Classification Using Byte-Level n-Gram Embeddings
Xiang Zhang, Alexandre Drouin, Raymond Li
This article introduces byteSteady -- a fast model for classification using byte-level n-gram embeddings. byteSteady assumes that each input comes as a sequence of bytes. A represe…
DoomArena: A framework for Testing AI Agents Against Evolving Security Threats
Leo Boisvert, Mihir Bansal, Chandra Kiran Reddy Evuru +9
We present DoomArena, a security evaluation framework for AI agents. DoomArena is designed on three principles: 1) It is a plug-in framework and integrates easily into realistic ag…
Benchmarking Bayesian Causal Discovery Methods for Downstream Treatment Effect Estimation
Chris Chinenye Emezue, Alexandre Drouin, Tristan Deleu +2
The practical utility of causality in decision-making is widespread and brought about by the intertwining of causal discovery and causal inference. Nevertheless, a notable gap exis…
In Search of Robust Measures of Generalization
Gintare Karolina Dziugaite, Alexandre Drouin, Brady Neal +5
One of the principal scientific challenges in deep learning is explaining generalization, i.e., why the particular way the community now trains networks to achieve small training e…
Toward Foundation Models for Earth Monitoring: Proposal for a Climate Change Benchmark
Alexandre Lacoste, Evan David Sherwin, Hannah Kerner +9
Recent progress in self-supervision shows that pre-training large neural networks on vast amounts of unsupervised data can lead to impressive increases in generalisation for downst…
InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation
Gaurav Sahu, Abhay Puri, Juan Rodriguez +11
Data analytics is essential for extracting valuable insights from data that can assist organizations in making effective decisions. We introduce InsightBench, a benchmark dataset w…
Dr-CiK: A Testbed for Foresight-Driven Agents
Yihong Tang, Andrew Robert Williams, Arjun Ashok +6
Time series forecasting in real-world settings often depends not only on historical observations, but also on external context that must be actively discovered from noisy, heteroge…
Backdoor Decontamination Dynamics in LLM Agents
Gabriel Huang, Abhay Puri, Léo Boisvert +4
Open-weight LLM agents are vulnerable to backdoors installed during fine-tuning, which may be undetectable if the trigger conditions are never met during testing. Assuming defender…
MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents
Alexander Gurung, Spandana Gella, Alexandre Drouin +3
Deep research agents increasingly combine private local documents with external tools like web retrieval, creating a privacy risk: an agent's external queries may leak sensitive in…
Invariant Causal Set Covering Machines
Thibaud Godon, Baptiste Bauvin, Pascal Germain +2
Rule-based models, such as decision trees, appeal to practitioners due to their interpretable nature. However, the learning algorithms that produce such models are often vulnerable…
Embedding Propagation: Smoother Manifold for Few-Shot Classification
Pau RodrÃguez, Issam Laradji, Alexandre Drouin +1
Few-shot classification is challenging because the data distribution of the training set can be widely different to the test set as their classes are disjoint. This distribution sh…
Context is Key: A Benchmark for Forecasting with Essential Textual Information
Andrew Robert Williams, Arjun Ashok, Ãtienne Marcotte +8
Forecasting is a critical task in decision-making across numerous domains. While historical numerical data provide a start, they fail to convey the complete context for reliable an…
Deep Learning for Electromyographic Hand Gesture Signal Classification Using Transfer Learning
Ulysse Côté-Allard, Cheikh Latyr Fall, Alexandre Drouin +5
In recent years, deep learning algorithms have become increasingly more prominent for their unparalleled ability to automatically learn discriminant features from large amounts of…
Capture the Flag: Uncovering Data Insights with Large Language Models
Issam Laradji, Perouz Taslakian, Sai Rajeswar +6
The extraction of a small number of relevant insights from vast amounts of data is a crucial component of data-driven decision-making. However, accomplishing this task requires con…
The BrowserGym Ecosystem for Web Agent Research
Thibault Le Sellier De Chezelles, Maxime Gasse, Alexandre Drouin +17
The BrowserGym ecosystem addresses the growing need for efficient evaluation and benchmarking of web agents, particularly those leveraging automation and Large Language Models (LLM…
Causal Representation Learning in Temporal Data via Single-Parent Decoding
Philippe Brouillard, Sébastien Lachapelle, Julia Kaltenborn +6
Scientific research often seeks to understand the causal structure underlying high-level variables in a system. For example, climate scientists study how phenomena, such as El Niñ…
Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain
Léo Boisvert, Léo Boisvert, Abhay Puri +8
While finetuning AI agents on interaction data -- such as web browsing or tool use -- improves their capabilities, it also introduces critical security vulnerabilities within the a…
RandomSCM: interpretable ensembles of sparse classifiers tailored for omics data
Thibaud Godon, Pier-Luc Plante, Baptiste Bauvin +3
Background: Understanding the relationship between the Omics and the phenotype is a central problem in precision medicine. The high dimensionality of metabolomics data challenges l…
WorkArena++: Towards Compositional Planning and Reasoning-based Common Knowledge Work Tasks
Léo Boisvert, Megh Thakkar, Maxime Gasse +6
The ability of large language models (LLMs) to mimic human-like intelligence has led to a surge in LLM-based autonomous agents. Though recent LLMs seem capable of planning and reas…
DRBench: A Realistic Benchmark for Enterprise Deep Research
Amirhossein Abaskohi, Tianyi Chen, Miguel Muñoz-Mármol +11
We introduce DRBench, a benchmark for evaluating AI agents on complex, open-ended deep research tasks in enterprise settings. Unlike prior benchmarks that focus on simple questions…