From the 1 of 2 linked papers with an AI index.
2 papers
cs.CV2026
OVEarth-Bench: Evaluating Category Breadth and Query Diversity for Open-Vocabulary Earth Observation
Kaiyu Li, Zepeng Xin, Zixuan Jiang +4
The paper presents OVEarth-Bench, a new benchmark for open-vocabulary Earth observation that evaluates models on a wide hierarchical set of categories and diverse query types, supp…
cs.AI2026
Act As a Real Researcher: A Suite of Benchmarks Evaluating Frontier LLMs and Agentic Harnesses in Research Lifecycle
Jiayu Wang, Weijiang Lv, Bowen Fu +8
As foundation models advance and agent scaffolding becomes increasingly sophisticated, agents have demonstrated remarkable proficiency in complex, long-horizon coding tasks and eve…