Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Self-Driving Datasets: From 20 Million Papers to Nuanced Biomedical Knowledge at Scale
Haydn Jones, Yimeng Zeng, Alden Rose +11
Manually curated biomedical repositories -- spanning bioactivity, genomics, and chemistry -- are expensive to maintain, lag behind primary literature, and discard experimental cont…
cs.LG2025
How and Why LLMs Generalize: A Fine-Grained Analysis of LLM Reasoning from Cognitive Behaviors to Low-Level Patterns
Haoyue Bai, Yiyou Sun, Wenjie Hu +5
Large Language Models (LLMs) display strikingly different generalization behaviors: supervised fine-tuning (SFT) often narrows capability, whereas reinforcement-learning (RL) tunin…
cs.LG2025
A Dataset for Distilling Knowledge Priors from Literature for Therapeutic Design
Haydn Thomas Jones, Natalie Maus, Josh Magnus Ludan +9
AI-driven discovery can greatly reduce design time and enhance new therapeutics' effectiveness. Models using simulators explore broad design spaces but risk violating implicit cons…