1.1k citations · 1.9k across the 31 of their papers we have counts for
30 papers
Supervised Contrastive Block Disentanglement
Taro Makino, Ji Won Park, Natasa Tagasovska +11
Real-world datasets often combine data collected under different experimental conditions. This yields larger datasets, but also introduces spurious correlations that make it diffic…
HERITAGE: An End-to-End Web Platform for Processing Korean Historical Documents in Hanja
Seyoung Song, Haneul Yoo, Jiho Jin +2
While Korean historical documents are invaluable cultural heritage, understanding those documents requires in-depth Hanja expertise. Hanja is an ancient language used in Korea befo…
Concept Bottleneck Language Models For protein design
Aya Abdelsalam Ismail, Tuomas Oikarinen, Amy Wang +8
We introduce Concept Bottleneck Protein Language Models (CB-pLM), a generative masked language model with a layer where each neuron corresponds to an interpretable concept. Our arc…
Antibody DomainBed: Out-of-Distribution Generalization in Therapeutic Protein Design
Nataša Tagasovska, Ji Won Park, Matthieu Kirchmeyer +8
Machine learning (ML) has demonstrated significant promise in accelerating drug design. Active ML-guided optimization of therapeutic molecules typically relies on a surrogate model…
Closed-Form Test Functions for Biophysical Sequence Optimization Algorithms
Samuel Stanton, Robert Alberstein, Nathan Frey +2
There is a growing body of work seeking to replicate the success of machine learning (ML) on domains like computer vision (CV) and natural language processing (NLP) to applications…
Following Length Constraints in Instructions
Weizhe Yuan, Ilia Kulikov, Ping Yu +4
Aligned instruction following models can better fulfill user requests than their unaligned counterparts. However, it has been shown that there is a length bias in evaluation of suc…