5 papers
Embedding Samples Dispatching for Recommendation Model Training in Edge Environments
Guopeng Li, Haisheng Tan, Chi Zhang +5
Training deep learning recommendation models (DLRMs) on edge workers brings several benefits, particularly in terms of data privacy protection, low latency and personalization. How…
A Plan Reuse Mechanism for LLM-Driven Agent
Guopeng Li, Ruiqi Wu, Haisheng Tan
Integrating large language models (LLMs) into personal assistants, like Xiao Ai and Blue Heart V, effectively enhances their ability to interact with humans, solve complex tasks, a…
Astraea: A State-Aware Scheduling Engine for LLM-Powered Agents
Hongqiu Ni, Jiabao Zhang, Guopeng Li +4
Large Language Models (LLMs) are increasingly being deployed as intelligent agents. Their multi-stage workflows, which alternate between local computation and calls to external net…
PICE: A Semantic-Driven Progressive Inference System for LLM Serving in Cloud-Edge Networks
Huiyou Zhan, Xuan Zhang, Haisheng Tan +4
Large language models (LLMs), while driving a new wave of interactive AI applications across numerous domains, suffer from high inference costs and heavy cloud dependency. Motivate…
Real-Time Neural-Enhancement for Online Cloud Gaming
Shan Jiang, Zhenhua Han, Haisheng Tan +6
Online Cloud gaming demands real-time, high-quality video transmission across variable wide-area networks (WANs). Neural-enhanced video transmission algorithms employing super-reso…