1 citations · 1 across the 11 of their papers we have counts for
Showing 2024Show all
3 papers · 1 filter
cs.LG2024
Ensuring Fair LLM Serving Amid Diverse Applications
Redwan Ibne Seraj Khan, Kunal Jain, Haiying Shen +12
In a multi-tenant large language model (LLM) serving platform hosting diverse applications, some users may submit an excessive number of requests, causing the service to become una…
cs.DC2024
PecSched: Preemptive and Efficient Cluster Scheduling for LLM Inference
Zeyu Zhang, Haiying Shen
The scaling of transformer-based Large Language Models (LLMs) has significantly expanded their context lengths, enabling applications where inputs exceed 100K tokens. Our analysis…
cs.LG2024
FDC: Fast KV Dimensionality Compression for Efficient LLM Inference
Zeyu Zhang, Haiying Shen
In large-language models, memory constraints in the Key-Value Cache (KVC) pose a challenge during inference. In this work, we propose FDC, a fast KV dimensionality compression syst…