3 papers
cs.AI2026
Planner-as-Router: Joint Plan-Time Model Routing for Cost-Efficient Multi-Agent Workflows
Vivek Kumar Singh, Preeti Priyam, Gautam Bhowmick
Running large language model (LLM) agents in production gets expensive fast. A frontier model (the largest, most capable tier) is accurate but can cost 25 times what a small model…
cs.AI2026
Total Cost of Agency: Exact Attribution of Memory Injection Cost in Multi-Agent LLM Workflows
Vivek Kumar Singh, Preeti Priyam, Gautam Bhowmick
Every node in a multi-agent large language model (LLM) workflow retrieves context from memory and injects it into its prompt, where those injected tokens are billed as input tokens…
cs.SE2026
ChurnBench: A Drift-Aware Benchmark Demonstrating That Refresh Scheduling, Not Cache Age, Governs Staleness in Agentic AI
Vivek Kumar Singh, Preeti Priyam, Gautam Bhowmick
In production, agentic systems answer questions over data that lives in several places and keeps changing: licenses are reassigned, users offboarded, prices changed, contracts rene…