works on

From the 1 of 18 linked papers with an AI index.

collaborators

18 papers

cs.CV2026

Chartography: A Benchmark for Professional Chart Understanding

Suhaas Garre, Chris Mutty, Sushant Mehta +1

Professionals across medicine, engineering, finance, manufacturing, and the sciences often make consequential decisions from charts. Existing chart benchmarks do not sufficiently m…

cs.AI2026

Post-Training on Office Work Improves Software Engineering: A Behavioral Account of Cross-Domain Transfer

Logan Ritchie, Sushant Mehta, Liudas Panavas +1

Long-horizon tasks require agents to maintain coherent state and goals across nested and branching work. We call this capability goal-directed execution (GDE): the repeated applica…

cs.SE2026

Cross-Benchmark Generalization in Long-Horizon Agents

Sushant Mehta, Logan Ritchie, Liudas Panavas +1

For reinforcement learning (RL) in self-contained environments, a policy can get rewards by exploiting environment-specific regularities (tool schemas, grader parsing, task templat…

cs.AI2026

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

Liudas Panavas, Sebastian Minus, Bradley Monton +4

Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to…

cs.CV2026

GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents

Suhaas Garre, Emily Ritchie, Sushant Mehta +1

The paper introduces GDP.pdf, a benchmark of professional PDF documents paired with realistic questions to evaluate grounded multimodal reasoning, and reports that current state‑of…

cs.AI2026

ComplexConstraints and Beyond: Expert Rubrics for RLVR

Sushant Mehta, Liudas Panavas, Suhaas Garre +1

Evaluation protocols can lag behind LLM capabilities. Programmatically verified benchmarks cover narrow surface constraints, whereas real-world instruction following and agentic wo…