activity
20182022
most citedProtein Representation Learning by Geometric Structure Pretraining

40 citations · 58 across the 4 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2022★ 40 cited

Protein Representation Learning by Geometric Structure Pretraining

Zuobai Zhang, Minghao Xu, Arian Jamasb +4

Learning effective protein representations is critical in a variety of tasks in biology such as predicting protein function or structure. Existing approaches usually pretrain prote…

cs.LG2021★ 14 cited

Fold2Seq: A Joint Sequence(1D)-Fold(3D) Embedding-based Generative Model for Protein Design

Yue Cao, Payel Das, Vijil Chenthamarakshan +3

Designing novel protein sequences for a desired 3D topological fold is a fundamental yet non-trivial task in protein engineering. Challenges exist due to the complex sequence--fold…

cs.LG2020

Optimizing Molecules using Efficient Queries from Property Evaluations

Samuel Hoffman, Vijil Chenthamarakshan, Kahini Wadhawan +2

Machine learning based methods have shown potential for optimizing existing molecules with more desirable properties, a critical step towards accelerating new chemical discovery. H…

cs.LG2020

Accelerating Antimicrobial Discovery with Controllable Deep Generative Models and Molecular Dynamics

Payel Das, Tom Sercu, Kahini Wadhawan +12

De novo therapeutic design is challenged by a vast chemical repertoire and multiple constraints, e.g., high broad-spectrum potency and low toxicity. We propose CLaSS (Controlled La…

cs.LG2020

CogMol: Target-Specific and Selective Drug Design for COVID-19 Using Deep Generative Models

Vijil Chenthamarakshan, Payel Das, Samuel C. Hoffman +8

The novel nature of SARS-CoV-2 calls for the development of efficient de novo drug design approaches. In this study, we propose an end-to-end framework, named CogMol (Controlled Ge…

cs.LG2019

A Sequential Set Generation Method for Predicting Set-Valued Outputs

Tian Gao, Jie Chen, Vijil Chenthamarakshan +1

Consider a general machine learning setting where the output is a set of labels or sequences. This output set is unordered and its size varies with the input. Whereas multi-label c…