4 papers
TENT: A Declarative Slice Spraying Engine for Performant and Resilient Data Movement in Disaggregated LLM Serving
Feng Ren, Ruoyu Qin, Teng Ma +16
Modern GPU clusters rely on complex, heterogeneous interconnects. As large language model (LLM) serving shifts toward agentic reasoning, KVCache becomes a first-class mobile asset,…
TrEnv-X: Transparently Share Serverless Execution Environments Across Different Functions and Nodes
Jialiang Huang, Teng Ma, Zheng Liu +11
Serverless computing is renowned for its computation elasticity, yet its full potential is often constrained by the requirement for functions to operate within local and dedicated…
Efficient Graph-Based Approximate Nearest Neighbor Search Achieving: Low Latency Without Throughput Loss
Jingjia Luo, Mingxing Zhang, Kang Chen +4
The increase in the dimensionality of neural embedding models has enhanced the accuracy of semantic search capabilities but also amplified the computational demands for Approximate…
Efficient Heterogeneous Large Language Model Decoding with Model-Attention Disaggregation
Shaoyuan Chen, Wencong Xiao, Yutong Lin +5
Transformer-based large language models (LLMs) exhibit impressive performance in generative tasks but also introduce significant challenges in real-world serving due to inefficient…