Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026★ 1 cited
LLM Novice Uplift on Dual-Use, In Silico Biology Tasks
Chen Bo Calvin Zhang, Christina Q. Knight, Nicholas Kruus +16
Large language models (LLMs) perform increasingly well on biology benchmarks, but it remains unclear whether they uplift novice users -- i.e., enable humans to perform better than…
cs.AI2025
Search-Time Data Contamination
Ziwen Han, Meher Mankikar, Julian Michael +1
Data contamination refers to the leakage of evaluation data into model training data, resulting in overfitting to supposedly held-out test sets and compromising test validity. We i…