most citedOpenProteinSet: Training data for structural biology at scale

15 citations · 19 across the 5 of their papers we have counts for

collaborators

5 papers

q-bio.BM20243 cited

Antibody DomainBed: Out-of-Distribution Generalization in Therapeutic Protein Design

Nataša Tagasovska, Ji Won Park, Matthieu Kirchmeyer +8

Machine learning (ML) has demonstrated significant promise in accelerating drug design. Active ML-guided optimization of therapeutic molecules typically relies on a surrogate model…

cs.LG2024

NEBULA: Neural Empirical Bayes Under Latent Representations for Efficient and Controllable Design of Molecular Libraries

Ewa M. Nowara, Pedro O. Pinheiro, Sai Pooja Mahajan +4

We present NEBULA, the first latent 3D generative model for scalable generation of large molecular libraries around a seed compound of interest. Such libraries are crucial for scie…

cs.LG2024

Closed-Form Test Functions for Biophysical Sequence Optimization Algorithms

Samuel Stanton, Robert Alberstein, Nathan Frey +2

There is a growing body of work seeking to replicate the success of machine learning (ML) on domains like computer vision (CV) and natural language processing (NLP) to applications…

q-bio.BM202315 cited

OpenProteinSet: Training data for structural biology at scale

Gustaf Ahdritz, Nazim Bouatta, Sachin Kadyan +7

Multiple sequence alignments (MSAs) of proteins encode rich biological information and have been workhorses in bioinformatic methods for tasks like protein design and protein struc…

cs.LG20231 cited

SupSiam: Non-contrastive Auxiliary Loss for Learning from Molecular Conformers

Michael Maser, Ji Won Park, Joshua Yao-Yu Lin +3

We investigate Siamese networks for learning related embeddings for augmented samples of molecular conformers. We find that a non-contrastive (positive-pair only) auxiliary task ai…