papers

Publications (8)

cs.LG2023

AVIDa-hIL6: A Large-Scale VHH Dataset Produced from an Immunized Alpaca for Predicting Antigen-Antibody Interactions

Hirofumi Tsuruta, Hiroyuki Yamazaki, Ryota Maeda +8

Antibodies have become an important class of therapeutic agents to treat human diseases. To accelerate therapeutic antibody discovery, computational methods, especially machine lea…

cs.LG2020

Population-Based Black-Box Optimization for Biological Sequence Design

Christof Angermueller, David Belanger, Andreea Gane +5

The use of black-box optimization for the design of new biological sequences is an emerging research area with potentially revolutionary impact. The cost and latency of wet-lab exp…

cs.LG2020

Masked Language Modeling for Proteins via Linearly Scalable Long-Context Transformers

Krzysztof Choromanski, Valerii Likhosherstov, David Dohan +8

Transformer models have achieved state-of-the-art results across a diverse range of domains. However, concern over the cost of training the attention mechanism to learn complex dep…

cs.LG2022

Rethinking Attention with Performers

Krzysztof Choromanski, Valerii Likhosherstov, David Dohan +10

We introduce Performers, Transformer architectures which can estimate regular (softmax) full-rank-attention Transformers with provable accuracy, but using only linear (as opposed t…

physics.chem-ph2019

Noisy, sparse, nonlinear: Navigating the Bermuda Triangle of physical inference with deep filtering

Carl Poelking, Yehia Amar, Alexei Lapkin +1

Capturing the microscopic interactions that determine molecular reactivity poses a challenge across the physical sciences. Even a basic understanding of the underlying reaction mec…

cs.LG2019

Using Attribution to Decode Dataset Bias in Neural Network Models for Chemistry

Kevin McCloskey, Ankur Taly, Federico Monti +2

Deep neural networks have achieved state of the art accuracy at classifying molecules with respect to whether they bind to specific protein targets. A key breakthrough would occur…

q-bio.BM2022

Meaningful machine learning models and machine-learned pharmacophores from fragment screening campaigns

Carl Poelking, Gianni Chessari, Christopher W. Murray +3

Machine learning (ML) is widely used in drug discovery to train models that predict protein-ligand binding. These models are of great value to medicinal chemists, in particular if…

q-bio.BM2020

Attribution Methods Reveal Flaws in Fingerprint-Based Virtual Screening

Vikram Sundar, Lucy Colwell

Fingerprint-based models for protein-ligand binding have demonstrated outstanding success on benchmark datasets; however, these models may not learn the correct binding rules. To a…