2 papers
cs.LG2026
WANDR: A Benchmark for Wide and Deep Research
Vitaliy Polshkov, Marcin Pitera, Jeremy Yang +7
WANDR (Wide ANd Deep Research) is a benchmark of 500 realistic, challenging data-collection tasks for research agents. Each task requires a system to discover a large set of entiti…
cs.CL2026
U-MATH: A University-Level Benchmark for Evaluating Mathematical Skills in LLMs
Konstantin Chernyshev, Vitaliy Polshkov, Ekaterina Artemova +4
The current evaluation of mathematical skills in LLMs is limited, as existing benchmarks are either relatively small, primarily focus on elementary and high-school problems, or lac…