4 papers
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
Weiye Wang, Chen Chen, Junxue Zhang +7
Distributed prefix caching has become a core technique for efficient LLM serving. However, for long-context requests with high cache hit ratios, retrieving reusable KVCache blocks…
Unlocking Full Efficiency of Token Filtering in Large Language Model Training
Di Chai, Pengbo Li, Feiyuan Zhang +7
Token filtering has been proposed to enhance the utility of large language models (LLMs) by eliminating inconsequential tokens during training. While usingfewer tokens is expected…
Swift: Rethinking RDMA Control Plane for Elastic Computing
Junxue Zhang, Han Tian, Xinyang Huang +5
Elastic computing enables dynamic scaling to meet workload demands, and Remote Direct Memory Access (RDMA) enhances this by providing high-throughput, low-latency network communica…
FLASH-FHE: A Heterogeneous Architecture for Fully Homomorphic Encryption Acceleration
Junxue Zhang, Xiaodian Cheng, Gang Cao +6
While many hardware accelerators have recently been proposed to address the inefficiency problem of fully homomorphic encryption (FHE) schemes, none of them is able to deliver opti…