2 papers
cs.AI2026
AgentIdeaBench: Benchmarking Scientific Ideation in the Agent Era
Yunxiang Mo, Tianshi Zheng, Yisen Gao +7
Scientific ideation is the capacity to formulate novel and testable hypotheses from scientific evidence, and autonomous AI scientists depend on it. Existing evaluations largely ass…
cs.AI2025
NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM Agents
Tianshi Zheng, Kelvin Kiu-Wai Tam, Newt Hue-Nam K. Nguyen +10
Large language models are emerging as powerful tools for scientific law discovery, a foundational challenge in AI-driven science. However, existing benchmarks for this task suffer…