3 papers
cs.DC2026
When Does Disaggregation Pay? Simulating Prefill--Decode--Attention--FFN Specialization for Agentic LLM Inference
Przemyslaw Forys, Haoran Wu, Can Xiao +8
Agentic inference now dominates the LLM inference landscape, requiring LLMs to actively engage in multi-turn interactions with tool-calling capabilities. This introduces a more com…
cs.AR2026
MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs
Haoran Wu, Zeyu Cao, Yao Lai +15
Emerging agentic LLM workloads are driving rapidly growing demand on both memory capacity and bandwidth, with different phases of inference (e.g., prefill and decode) imposing dist…
cs.AI2026
Heterogeneous Computing: The Key to Powering the Future of AI Agent Inference
Yiren Zhao, Junyi Liu
AI agent inference is driving an inference heavy datacenter future and exposes bottlenecks beyond compute - especially memory capacity, memory bandwidth and high-speed interconnect…