6 papers
Beyond Greedy Chunking: SLO-Aware Sliding-Window Scheduling for LLM Inference
Yuansheng Chen, Yue Zhang, Xuan Mo +2
With the rapid growth of interactive applications in large language model (LLM) online services, maintaining high system throughput while ensuring user-perceived latency has become…
Analysis of Asynchronous Federated Learning: Unraveling the Interactions between Gradient Compression, Delay, and Data Heterogeneity
Diying Yang, Yingwei Hou, Weigang Wu
In practical federated learning (FL), the large communication overhead between clients and the server is often a significant bottleneck. Gradient compression methods can effectivel…
MaaSO: SLO-aware Orchestration of Heterogeneous Model Instances for MaaS
Mo Xuan, Zhang yue, Wu Weigang
Model-as-a-Service (MaaS) platforms face diverse Service Level Objective (SLO) requirements stemming from various large language model (LLM) applications, manifested in contextual…
Cool-Fusion: Fuse Large Language Models without Training
Cong Liu, Xiaojun Quan, Yan Pan +3
We focus on the problem of fusing two or more heterogeneous large language models (LLMs) to leverage their complementary strengths. One of the challenges of model fusion is high co…
Chain of Methodologies: Scaling Test Time Computation without Training
Cong Liu, Jie Wu, Weigang Wu +3
Large Language Models (LLMs) often struggle with complex reasoning tasks due to insufficient in-depth insights in their training data, which are typically absent in publicly availa…
VersiCode: Towards Version-controllable Code Generation
Tongtong Wu, Weigang Wu, Xingyu Wang +7
Large Language Models (LLMs) have made tremendous strides in code generation, but existing research fails to account for the dynamic nature of software development, marked by frequ…