collaborators

7 papers

cs.LG2026

SAKI: Score-Aware Low-Rank Key Indexing with Random-Matrix Noise Correction for KV Retrieval

Lin Zhang

Existing low rank KV cache methods preserve either model weights or key variance, neither of which directly reflects the attention scores used during inference. We derive the expec…

cs.DC2026

Zellige: Moldable Sequence Placement for Mixed Image-Video DiT Training

Guangyu Xiang, Xueze Kang, Minwei Zhao +4

High-quality video generation requires training Diffusion Transformers (DiTs) jointly on image and video data, posing a mixed-length sequence training problem across GPUs. Existing…

cs.DC2026

Xema: Efficient Diffusion Serving through Fine-Grained Memory Management and Auto-Configuration

Xueze Kang, Guangyu Xiang, Suyi Li +4

Xema is a system that reduces GPU memory usage for diffusion model serving by analyzing tensor lifetimes to apply targeted memory mitigation and by planning parallelism and concurr…

cs.DC2026

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding

Guangyu Xiang, Xueze Kang, Lin Zhang +4

LLM serving is increasingly dominated by long and dynamic decode workloads from agents, reasoning models, and extended conversations. When bursty long-context demand exceeds deploy…

cs.CL2026

HyLRA: Hybrid Layer Reuse Attention for Efficient Long-Context Inference

Xuan Ai, Qingqing Yang, Peng Wang +4

Long-context inference in Large Language Models (LLMs) is bottlenecked by the quadratic computation complexity of attention and the substantial memory footprint of Key-Value (KV) c…

cs.DC2025

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap

Wenxiang Lin, Xinglin Pan, Lin Zhang +3

The sparsely activated mixture-of-experts (MoE) transformer has become a common architecture for large language models (LLMs) due to its sparsity, which requires fewer computationa…