1 paper · 1 filter
Sheng Liang, Yongyue Zhang, Nathanael Brian +4
Agentic LLM pipelines face escalating inference costs as context accumulates across retrieval, tool use, and multi-turn interactions. To control latency, deployments routinely comp…