2 papers
cs.AI2026
Nexus: Depth-Adaptive KV-Cache Splicing and Retrieval-Decoupled Tool Routing for Agentic LLMs on Unified Memory
Mustafa Arslan
Agentic large language models (LLMs) on the Model Context Protocol (MCP) re-encode verbose tool schemas every turn, so prefill - quadratic in sequence length - dominates time-to-fi…
cs.AI2026
Aeon: High-Performance Neuro-Symbolic Memory Management for Long-Horizon LLM Agents
Mustafa Arslan
Large Language Models (LLMs) are fundamentally constrained by the quadratic computational cost of self-attention and the "Lost in the Middle" phenomenon, where reasoning capabiliti…