3 papers
cs.PF2026
Compass: Dissecting Communication and Computation Operators for Efficient LLM Training
Guangyu Xiang, Lin Zhang, Haoxuan Yu +3
Overlapping communication and computation operators is a common practice to hide communication overheads, accelerating large language models (LLMs) training on GPU clusters. Existi…
cs.CV2026
AdaFocus: Adaptive Relevance-Diversity Sampling with Zero-Cache Look-back for Efficient Long Video Understanding
Xiao Yang, Yingzhe Ma, Haoxuan Yu +2
Long video understanding is heavily bottlenecked by a rigid one-shot paradigm: existing methods either densely encode videos at prohibitive memory and latency costs, or aggressivel…
cs.DC2024
SGDRC: Software-Defined Dynamic Resource Control for Concurrent DNN Inference on NVIDIA GPUs
Yongkang Zhang, Haoxuan Yu, Chenxia Han +7
Cloud service providers heavily colocate high-priority, latency-sensitive (LS), and low-priority, best-effort (BE) DNN inference services on the same GPU to improve resource utiliz…