10 papers
Self-Driving Datasets: From 20 Million Papers to Nuanced Biomedical Knowledge at Scale
Haydn Jones, Yimeng Zeng, Alden Rose +11
Manually curated biomedical repositories -- spanning bioactivity, genomics, and chemistry -- are expensive to maintain, lag behind primary literature, and discard experimental cont…
Purely Agent-Driven Black-Box Optimization for Biological Design
Natalie Maus, Yimeng Zeng, Haydn Thomas Jones +11
Many key challenges in biological design -- such as small-molecule drug discovery, antimicrobial peptide development, and protein engineering -- can be framed as black-box optimiza…
Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference Models
Anirudh Bharadwaj, Chaitanya Malaviya, Nitish Joshi +1
Language models serve as proxies for human preference judgements in alignment and evaluation, yet they exhibit systematic miscalibration, prioritizing superficial patterns over sub…
ResearchQA: Evaluating Scholarly Question Answering at Scale Across 75 Fields with Survey-Mined Questions and Rubrics
Li S. Yifei, Allen Chang, Chaitanya Malaviya +1
Evaluating long-form responses to research queries heavily relies on expert annotators, restricting attention to areas like AI where researchers can conveniently enlist colleagues.…
Contra4: Evaluating Contrastive Cross-Modal Reasoning in Audio, Video, Image, and 3D
Artemis Panagopoulou, Le Xue, Honglu Zhou +6
Real-world decision-making often begins with identifying which modality contains the most relevant information for a given query. While recent multimodal models have made impressiv…
A Dataset for Distilling Knowledge Priors from Literature for Therapeutic Design
Haydn Thomas Jones, Natalie Maus, Josh Magnus Ludan +9
AI-driven discovery can greatly reduce design time and enhance new therapeutics' effectiveness. Models using simulators explore broad design spaces but risk violating implicit cons…