activity
20242026
collaborators

12 papers

cs.LG2026

Self-Driving Datasets: From 20 Million Papers to Nuanced Biomedical Knowledge at Scale

Haydn Jones, Yimeng Zeng, Alden Rose +11

Manually curated biomedical repositories -- spanning bioactivity, genomics, and chemistry -- are expensive to maintain, lag behind primary literature, and discard experimental cont…

cs.LG2026

Purely Agent-Driven Black-Box Optimization for Biological Design

Natalie Maus, Yimeng Zeng, Haydn Thomas Jones +11

Many key challenges in biological design -- such as small-molecule drug discovery, antimicrobial peptide development, and protein engineering -- can be framed as black-box optimiza…

cs.CL2026

Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference Models

Anirudh Bharadwaj, Chaitanya Malaviya, Nitish Joshi +1

Language models serve as proxies for human preference judgements in alignment and evaluation, yet they exhibit systematic miscalibration, prioritizing superficial patterns over sub…

cs.CL2025

ResearchQA: Evaluating Scholarly Question Answering at Scale Across 75 Fields with Survey-Mined Questions and Rubrics

Li S. Yifei, Allen Chang, Chaitanya Malaviya +1

Evaluating long-form responses to research queries heavily relies on expert annotators, restricting attention to areas like AI where researchers can conveniently enlist colleagues.…

cs.AI2025

Contra4: Evaluating Contrastive Cross-Modal Reasoning in Audio, Video, Image, and 3D

Artemis Panagopoulou, Le Xue, Honglu Zhou +6

Real-world decision-making often begins with identifying which modality contains the most relevant information for a given query. While recent multimodal models have made impressiv…

cs.LG2025

A Dataset for Distilling Knowledge Priors from Literature for Therapeutic Design

Haydn Thomas Jones, Natalie Maus, Josh Magnus Ludan +9

AI-driven discovery can greatly reduce design time and enhance new therapeutics' effectiveness. Models using simulators explore broad design spaces but risk violating implicit cons…