Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation
Yunting Song, Matthew Watson, Peter Grabowski +1
The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality. As datasets scale, massive preference and instruction-tuning corpora inevitably accumula…
cs.LG2025
Assay2Mol: large language model-based drug design using BioAssay context
Yifan Deng, Spencer S. Ericksen, Anthony Gitter
Scientific databases aggregate vast amounts of quantitative data alongside descriptive text. In biochemistry, molecule screening assays evaluate candidate molecules' functional res…