From the 1 of 6 linked papers with an AI index.
6 papers
Chartography: A Benchmark for Professional Chart Understanding
Suhaas Garre, Chris Mutty, Sushant Mehta +1
Professionals across medicine, engineering, finance, manufacturing, and the sciences often make consequential decisions from charts. Existing chart benchmarks do not sufficiently m…
HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
Liudas Panavas, Sebastian Minus, Bradley Monton +4
Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to…
GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents
Suhaas Garre, Emily Ritchie, Sushant Mehta +1
The paper introduces GDP.pdf, a benchmark of professional PDF documents paired with realistic questions to evaluate grounded multimodal reasoning, and reports that current state‑of…
ComplexConstraints and Beyond: Expert Rubrics for RLVR
Sushant Mehta, Liudas Panavas, Suhaas Garre +1
Evaluation protocols can lag behind LLM capabilities. Programmatically verified benchmarks cover narrow surface constraints, whereas real-world instruction following and agentic wo…
Riemann-Bench: A Benchmark for Moonshot Mathematics
Suhaas Garre, Erik Knutsen, Sushant Mehta +1
Recent AI systems have achieved gold-medal-level performance on the International Mathematical Olympiad, demonstrating remarkable proficiency at competition-style problem solving.…
EnterpriseBench Corecraft: Training Generalizable Agents on High-Fidelity RL Environments
Sushant Mehta, Logan Ritchie, Suhaas Garre +3
We show that training AI agents on high-fidelity reinforcement learning environments produces capabilities that generalize beyond the training distribution. We introduce CoreCraft,…