2 papers
cs.AI2026
Divergent-Convergent Reasoning: Scaling Test-Time Compute through Structured Solution Synthesis
Bo Wen, Yuhao Chen, Erhan Bilal +3
Test-time compute can substantially improve Large Language Model (LLM) reasoning performance, yet how and when additional compute helps remains poorly understood. We study Divergen…
cs.AI2026
Learning Agent Execution for KV-Cache Management in Agentic Serving
Rui Zhang, Chaeeun Kim, Shaoting Feng +6
Multi-agent LLM systems have emerged as an important deployment paradigm for AI services, where each user request is decomposed into a sequence of specialized agents. Across these…