papers

Publications (10)

cs.LG2025

Invariant Causal Set Covering Machines

Thibaud Godon, Baptiste Bauvin, Pascal Germain +2

Rule-based models, such as decision trees, appeal to practitioners due to their interpretable nature. However, the learning algorithms that produce such models are often vulnerable…

q-bio.GN2014

Learning interpretable models of phenotypes from whole genome sequences with the Set Covering Machine

Alexandre Drouin, Sébastien Giguère, Vladana Sagatovich +4

The increased affordability of whole genome sequencing has motivated its use for phenotypic studies. We address the problem of learning interpretable models for discrete phenotypes…

q-bio.GN2013

Assemblathon 2: evaluating de novo methods of genome assembly in three vertebrate species

Keith R. Bradnam, Joseph N. Fass, Anton Alexandrov +88

Background - The process of generating raw genome sequence data continues to become cheaper, faster, and more accurate. However, assembly of such data into high-quality, finished g…

cs.LG2010

Feature Selection with Conjunctions of Decision Stumps and Learning from Microarray Data

Mohak Shah, Mario Marchand, Jacques Corbeil

One of the objectives of designing feature selection learning algorithms is to obtain classifiers that depend on a small number of attributes and have verifiable future performance…

q-bio.QM2026

Refnd: Preventing Data Leakage in Relational Datasets

Anthony Lavertu, Jacob Cote, Jacques Corbeil +2

Machine learning models trained on biochemical data are routinely evaluated using splits that fail to account for relational structure, causing information leakage and over-optimis…

q-bio.QM2014

Improved design and screening of high bioactivity peptides for drug discovery

Sébastien Giguère, François Laviolette, Mario Marchand +4

The discovery of peptides having high biological activity is very challenging mainly because there is an enormous diversity of compounds and only a minority have the desired proper…