most citedTowards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization

2 citations · 2 across the 2 of their papers we have counts for

collaborators

5 papers

cs.LG20262 cited

Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization

Jiantong Jiang, Peiyu Yang, Rui Zhang +1

Despite the rapid advancements of large language models (LLMs), LLM serving systems remain memory-intensive and costly. The key-value (KV) cache, which stores KV tensors during aut…

cs.DC2026

Coordinated Scheduling for MoE LLM Serving

Yifan Sun, Zhexiang Zhang, Jiantong Jiang +5

Serving Mixture-of-Experts (MoE) large language models (LLMs) is challenging because dynamic request workloads interact with sparse expert routing, creating both data-parallel (DP)…

cs.LG2026

Attribution-Guided Model Rectification of Unreliable Neural Network Behaviors

Peiyu Yang, Naveed Akhtar, Jiantong Jiang +1

The performance of neural network models deteriorates due to their unreliable behavior on non-robust features of corrupted samples. Owing to their opaque nature, rectifying models…

cs.DC2025

FLASH Viterbi: Fast and Adaptive Viterbi Decoding for Modern Data Systems

Ziheng Deng, Xue Liu, Jiantong Jiang +3

The Viterbi algorithm is a key operator for structured sequence inference in modern data systems, with applications in trajectory analysis, online recommendation, and speech recogn…

cs.CR2025

A Backdoor-based Explainable AI Benchmark for High Fidelity Evaluation of Attributions

Peiyu Yang, Naveed Akhtar, Jiantong Jiang +1

Attribution methods compute importance scores for input features to explain model predictions. However, assessing the faithfulness of these methods remains challenging due to the a…