1 citations · 1 across the 2 of their papers we have counts for
4 papers
Evaluating Large Language Models in Scientific Discovery
Zhangde Song, Jieyu Lu, Yuanqi Du +53
Large language models (LLMs) are increasingly applied to scientific research, yet prevailing science benchmarks probe decontextualized knowledge and overlook the iterative reasonin…
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
Xinming Tu, Tianze Wang, Yingzhou +4
As benchmarks grow in complexity, many apparent agent failures are not failures of the agent at all - they are failures of the benchmark itself: broken specifications, implicit ass…
CRISPR-GPT for Agentic Automation of Gene-editing Experiments
Yuanhao Qu, Kaixuan Huang, Ming Yin +11
The introduction of genome engineering technology has transformed biomedical research, making it possible to make precise changes to genetic information. However, creating an effic…
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning
Ming Yin, Yuanhao Qu, Ling Yang +2
We investigate how to teach large language models (LLMs) to perform scientific reasoning by leveraging expert discussions as a learning signal. Focusing on the genomics domain, we…