2 papers
cs.AI2026
Read Less, Solve More: Token-Efficient Sparse Reading for AI Agents
Zedong Liu, Jiaan Wu, Xinyang Ma +5
Long-horizon agents increasingly rely on repeated access to external artifacts, yet current reading interfaces often expose entire objects even when only sparse evidence is needed.…
cs.DC2026
KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
Zedong Liu, Xinyang Ma, Dejun Luo +9
LLMs are widely adopted in production, pushing inference systems to their limits. Disaggregated LLM serving (e.g., PD separation and KV state disaggregation) improves scalability a…