4 papers · 1 filter
CONCORD: Asynchronous Sparse Aggregation for Device-Cloud RAG under Document Isolation
Xuedong Hu, Zhiqing Tang, Zhi Yao +2
Retrieval-augmented generation (RAG) has emerged as a pivotal technique for improving language models by incorporating external knowledge at inference time. As device-cloud collabo…
HeraSys: Collaborative Serving of Multiple LLM Workflows via Fine-Grained End-to-End Optimization
Size Li, Zhiqing Tang, Hongrui Liang +4
The proliferation of Large Language Models (LLMs) has shifted serving systems from processing isolated requests to orchestrating high-concurrency, multi-tenant agentic workflows. H…
Adaptive AI Agent Placement and Migration in Edge Intelligence Systems
Xingdan Wang, Jiayi He, Zhiqing Tang +5
The rise of LLMs such as ChatGPT and Claude fuels the need for AI agents capable of real-time task handling. However, migrating data-intensive, multi-modal edge workloads to cloud…
Empowering Edge Intelligence: A Comprehensive Survey on On-Device AI Models
Xubin Wang, Zhiqing Tang, Jianxiong Guo +4
The rapid advancement of artificial intelligence (AI) technologies has led to an increasing deployment of AI models on edge and terminal devices, driven by the proliferation of the…