4 papers
ReasonOps: Operator Segmentation for LLM Reasoning Traces
Daniel Lee, Owen Queen, James Zou
Chain-of-thought traces from large reasoning models can span tens of thousands of tokens, yet we lack a vocabulary for describing their internal structure. Previous methods develop…
DSGym: A Holistic Framework for Evaluating and Training Data Science Agents
Fan Nie, Junlin Wang, Harper Hua +6
Data science agents promise to accelerate discovery and insight-generation by turning data into executable analyses and findings. Yet existing data science benchmarks fall short du…
Exploring the use of AI authors and reviewers at Agents4Science
Federico Bianchi, Owen Queen, Nitya Thakkar +2
There is growing interest in using AI agents for scientific research, yet fundamental questions remain about their capabilities as scientists and reviewers. To explore these questi…
CGBench: Benchmarking Language Model Scientific Reasoning for Clinical Genetics Research
Owen Queen, Harrison G. Zhang, James Zou
Variant and gene interpretation are fundamental to personalized medicine and translational biomedicine. However, traditional approaches are manual and labor-intensive. Generative l…