activity
20242026
most citedPAT: Accelerating LLM Decoding via Prefix-Aware Attention with Resource Efficient Multi-Tile Kernel

1 citations · 1 across the 3 of their papers we have counts for

collaborators

5 papers

cs.DC20261 cited

PAT: Accelerating LLM Decoding via Prefix-Aware Attention with Resource Efficient Multi-Tile Kernel

Jinjun Yi, Zhixin Zhao, Yitao Hu +7

LLM serving is increasingly dominated by decode attention, which is a memory-bound operation due to massive KV cache loading from global memory. Meanwhile, real-world workloads exh…

cs.DC2026

PARD: Enhancing Goodput for Inference Pipeline via Proactive Request Dropping

Zhixin Zhao, Yitao Hu, Simin Chen +8

Modern deep neural network (DNN) applications integrate multiple DNN models into inference pipelines with stringent latency requirements for customized tasks. To mitigate extensive…

cs.LG2026

Mosaic: Unlocking Long-Context Inference for Diffusion LLMs via Global Memory Planning and Dynamic Peak Taming

Liang Zheng, Bowen Shi, Yitao Hu +5

Diffusion-based large language models (dLLMs) have emerged as a promising paradigm, utilizing simultaneous denoising to enable global planning and iterative refinement. While these…

cs.DC2025

EPARA: Parallelizing Categorized AI Inference in Edge Clouds

Yubo Wang, Yubo Cui, Tuo Shi +5

With the increasing adoption of AI applications such as large language models and computer vision AI, the computational demands on AI inference systems are continuously rising, mak…

cs.DC2024

Harpagon: Minimizing DNN Serving Cost via Efficient Dispatching, Scheduling and Splitting

Zhixin Zhao, Yitao Hu, Ziqi Gong +5

Advances in deep neural networks (DNNs) have significantly contributed to the development of real-time video processing applications. Efficient scheduling of DNN workloads in cloud…