7 papers
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…
Benchmarking AI Agents for Addressing Scientific Challenges Across Scales
Tianyu Liu, Allen Xin Wang, Antonia Panescu +30
AI agents are increasingly being developed to accelerate scientific discovery, yet their practical capabilities in real research settings remain poorly understood. Existing benchma…
Towards Diverse Scientific Hypothesis Search with Large Language Models
Haorui Wang, Parshin Shojaee, Kazem Meidani +7
Large language models (LLMs) are on the rise for accelerating scientific discovery, most recently in advanced tasks such as generating valid scientific hypotheses. Yet in many disc…
Vision-Language Guided Hyperspectral Object Tracking via Semantics Fusion and Contextual Template Updating
Rui Yao, Yuhong Zhang, Kunyang Sun +4
Hyperspectral object tracking (HOT) leverages the rich spectral information provided by hyperspectral videos (HSVs), offering substantial potential for object tracking. However, ef…
Accelerating Scientific Discovery with Autonomous Goal-evolving Agents
Yuanqi Du, Botao Yu, Tianyu Liu +25
There has been unprecedented interest in developing agents that expand the boundary of scientific discovery, primarily by optimizing quantitative objective functions specified by s…
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…