3 papers
cs.LG2026
Efficient Serving for Dynamic Agent Workflows with Prediction-based KV-Cache Management
Haoyu Zheng, Fangcheng Fu, Jia Wu +6
LLM-based workflows compose specialized agents to execute complex tasks, and these agents usually share substantial context, allowing KV-Cache reuse to save computation. Existing a…
cs.AI2026
HEDP: A Hybrid Energy-Distance Prompt-based Framework for Domain Incremental Learning
Yu Feng, Zhen Tian, Haoran Luo +8
Domain Incremental Learning is a critical scenario that requires models to continuously adapt to new data domains without retraining. However, domain shifts often cause severe perf…
cs.DC2025
DOLMA: A Data Object Level Memory Disaggregation Framework for HPC Applications
Haoyu Zheng, Shouwei Gao, Jie Ren +1
Memory disaggregation is promising to scale memory capacity and improves utilization in HPC systems. However, the performance overhead of accessing remote memory poses a significan…