collaborators

7 papers

cs.DC2026

Odin: Primitive-Level Synchronization for Distributed Point-Based Neural Rendering

Zhenxiang Ma, Zeyu He, Yuanzhen Zhou +6

Point-based neural rendering (PBNR) represents 3D scenes as explicit, trainable primitives and underpins high-quality reconstruction and emerging embodied AI and world-model pipeli…

cs.CV2026

DiTango: Cost-Effective Parallel Diffusion Generation with Selective Attention State Reuse

Yuyang Chen, Runxin Zhong, Zan Zong +3

Recent advances in AI-generated content have driven widespread adoption of Diffusion Transformers (DiTs) for high-resolution, long-duration content generation. While parallelizatio…

cs.DC2025

FlashFuser: Expanding the Scale of Kernel Fusion for Compute-Intensive Operators via Inter-Core Connection

Ziyu Huang, Yangjie Zhou, Zihan Liu +8

The scaling of computation throughput continues to outpace improvements in memory bandwidth, making many deep learning workloads memory-bound. Kernel fusion is a key technique to a…

cs.LG2025

Lethe: Layer- and Time-Adaptive KV Cache Pruning for Reasoning-Intensive LLM Serving

Hui Zeng, Daming Zhao, Pengfei Yang +5

Generative reasoning with large language models (LLMs) often involves long decoding sequences, leading to substantial memory and latency overheads from accumulating key-value (KV)…

cs.LG2025

SpecRouter: Adaptive Routing for Multi-Level Speculative Decoding in Large Language Models

Hang Wu, Jianian Zhu, Yinghui Li +3

Large Language Models (LLMs) present a critical trade-off between inference quality and computational cost: larger models offer superior capabilities but incur significant latency,…

cs.DC2025

Jenga: Effective Memory Management for Serving LLM with Heterogeneity

Chen Zhang, Kuntai Du, Shu Liu +10

Large language models (LLMs) are widely used but expensive to run, especially as inference workloads grow. To lower costs, maximizing the request batch size by managing GPU memory…