works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

cs.CV2026

Chartography: A Benchmark for Professional Chart Understanding

Suhaas Garre, Chris Mutty, Sushant Mehta +1

Professionals across medicine, engineering, finance, manufacturing, and the sciences often make consequential decisions from charts. Existing chart benchmarks do not sufficiently m…

cs.AI2026

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

Liudas Panavas, Sebastian Minus, Bradley Monton +4

Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to…

cs.CV2026

GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents

Suhaas Garre, Emily Ritchie, Sushant Mehta +1

The paper introduces GDP.pdf, a benchmark of professional PDF documents paired with realistic questions to evaluate grounded multimodal reasoning, and reports that current state‑of…

cs.AI2026

ComplexConstraints and Beyond: Expert Rubrics for RLVR

Sushant Mehta, Liudas Panavas, Suhaas Garre +1

Evaluation protocols can lag behind LLM capabilities. Programmatically verified benchmarks cover narrow surface constraints, whereas real-world instruction following and agentic wo…

cs.AI2026

Riemann-Bench: A Benchmark for Moonshot Mathematics

Suhaas Garre, Erik Knutsen, Sushant Mehta +1

Recent AI systems have achieved gold-medal-level performance on the International Mathematical Olympiad, demonstrating remarkable proficiency at competition-style problem solving.…

cs.AI2026

EnterpriseBench Corecraft: Training Generalizable Agents on High-Fidelity RL Environments

Sushant Mehta, Logan Ritchie, Suhaas Garre +3

We show that training AI agents on high-fidelity reinforcement learning environments produces capabilities that generalize beyond the training distribution. We introduce CoreCraft,…