Publications (10)
Invariant Causal Set Covering Machines
Thibaud Godon, Baptiste Bauvin, Pascal Germain +2
Rule-based models, such as decision trees, appeal to practitioners due to their interpretable nature. However, the learning algorithms that produce such models are often vulnerable…
Learning interpretable models of phenotypes from whole genome sequences with the Set Covering Machine
Alexandre Drouin, Sébastien Giguère, Vladana Sagatovich +4
The increased affordability of whole genome sequencing has motivated its use for phenotypic studies. We address the problem of learning interpretable models for discrete phenotypes…
Assemblathon 2: evaluating de novo methods of genome assembly in three vertebrate species
Keith R. Bradnam, Joseph N. Fass, Anton Alexandrov +88
Background - The process of generating raw genome sequence data continues to become cheaper, faster, and more accurate. However, assembly of such data into high-quality, finished g…
Feature Selection with Conjunctions of Decision Stumps and Learning from Microarray Data
Mohak Shah, Mario Marchand, Jacques Corbeil
One of the objectives of designing feature selection learning algorithms is to obtain classifiers that depend on a small number of attributes and have verifiable future performance…
Refnd: Preventing Data Leakage in Relational Datasets
Anthony Lavertu, Jacob Cote, Jacques Corbeil +2
Machine learning models trained on biochemical data are routinely evaluated using splits that fail to account for relational structure, causing information leakage and over-optimis…
Improved design and screening of high bioactivity peptides for drug discovery
Sébastien Giguère, François Laviolette, Mario Marchand +4
The discovery of peptides having high biological activity is very challenging mainly because there is an enormous diversity of compounds and only a minority have the desired proper…