1 paper
Jiayi Qian, Yichong Zhang, Zishen Wan +4
Efficient serving of agentic workflows requires selecting each LLM node's model, verification policy, and backend to balance output quality, latency, and throughput. These assignme…