#gpu scheduling
try —
2 papers match
cs.DC2026
FlashDiff: Efficient Regional Execution and Scheduling for Diffusion Model Serving
Yaqi Qiao, Ping He, Songrun Xie +4
FlashDiff is a system that speeds up diffusion model inference by dynamically selecting which latent regions need further processing and efficiently scheduling those regions across…
#diffusion models#model serving#regional execution#GPU scheduling
cs.PF2026
EMO: Energy Efficiency Modeling and Optimization for AI Workloads
Jiyu Luo, Shaoyu Chen, Jingwei Sun +3
EMO is a lightweight framework that models and optimizes the energy consumption of GPU-accelerated AI workloads by detecting fine‑grained slack in asynchronous execution and applyi…
#energy efficiency#gpu scheduling#ai workloads#asynchronous execution