2 citations · 2 across the 2 of their papers we have counts for
5 papers
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization
Jiantong Jiang, Peiyu Yang, Rui Zhang +1
Despite the rapid advancements of large language models (LLMs), LLM serving systems remain memory-intensive and costly. The key-value (KV) cache, which stores KV tensors during aut…
Coordinated Scheduling for MoE LLM Serving
Yifan Sun, Zhexiang Zhang, Jiantong Jiang +5
Serving Mixture-of-Experts (MoE) large language models (LLMs) is challenging because dynamic request workloads interact with sparse expert routing, creating both data-parallel (DP)…
Attribution-Guided Model Rectification of Unreliable Neural Network Behaviors
Peiyu Yang, Naveed Akhtar, Jiantong Jiang +1
The performance of neural network models deteriorates due to their unreliable behavior on non-robust features of corrupted samples. Owing to their opaque nature, rectifying models…
FLASH Viterbi: Fast and Adaptive Viterbi Decoding for Modern Data Systems
Ziheng Deng, Xue Liu, Jiantong Jiang +3
The Viterbi algorithm is a key operator for structured sequence inference in modern data systems, with applications in trajectory analysis, online recommendation, and speech recogn…
A Backdoor-based Explainable AI Benchmark for High Fidelity Evaluation of Attributions
Peiyu Yang, Naveed Akhtar, Jiantong Jiang +1
Attribution methods compute importance scores for input features to explain model predictions. However, assessing the faithfulness of these methods remains challenging due to the a…