3 papers
cs.LG2026
AdvPlan-Bench: Adversarial Evaluation of Structured Plan-Generation Agents
Alina Kapanova, Arun Kanhai, Natan Vidra +1
Structured plan-generation agents are often evaluated as if a plan has quality in isolation, yet many realistic planning tasks require asking how a candidate behaves when another a…
cs.LG2026
MetaRoute-Bench: Evaluating Meta-Decision Policies for Agentic Workflow Routing
Natan Vidra, Alina Kapanova, Arun Kanhai +1
Agentic systems must repeatedly decide whether to answer directly, decompose a task, invoke a tool, execute code, delegate to a specialist, verify an intermediate result, or recove…
cs.LG2026
Learning Compositional Meta-Routing for Agentic Workflows: An Executable Benchmark
Natan Vidra, Alina Kapanova, Arun Kanhai +1
Agentic systems must decide not only what answer to produce, but which reasoning and execution operations should precede it. A controller may answer directly, decompose a request,…