Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Deployment-Time Memorization in Foundation-Model Agents
Lei, Chen, Guilin Zhang +8
Foundation-model agents are increasingly long-lived systems that remember users across interactions, making memorization an explicit deployment-time function rather than solely a p…
cs.AI2026
From Rigid to Dynamic: Entropy-Guided Adaptive Inference for Long-Context LLMs
Zhanchao Xu, Haoyang Li, Qingfa Xiao +4
Existing sparse attention and KV cache compression methods for long-context LLM inference typically apply fixed sparsity patterns or uniform budgets across all attention heads, ove…