2 papers
cs.AI2026
ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds
Ming Zhang, Zhenghao Xiang, Peizhong Gao +17
Scientific discovery begins where known problems end. There, AI systems must engage in exploration: framing hypotheses, designing experiments, and iterating on the results. However…
cs.AI2026
NovGauge: A Fine-Grained Benchmark for Diagnosing LLMs' Capability in Paper Novelty Assessment
Guoqiang Zhang, Kexin Tan, Ming Zhang +12
Large language models (LLMs) are increasingly used in peer review at major AI conferences, yet novelty remains a persistent weak point. Existing benchmarks assess novelty as a sing…