activity
20172026
most citedLRC: Dependency-Aware Cache Management for Data Analytics Clusters

5 citations · 7 across the 11 of their papers we have counts for

collaborators
Showing cs.DCShow all

7 papers · 1 filter

cs.DC2025

RollMux: Phase-Level Multiplexing for Disaggregated RL Post-Training

Tianyuan Wu, Lunxi Cao, Yining Wei +11

Rollout-training disaggregation is emerging as the standard architecture for Reinforcement Learning (RL) post-training, where memory-bound rollout and compute-bound training are ph…

cs.DC2025

InstGenIE: Generative Image Editing Made Efficient with Mask-aware Caching and Scheduling

Xiaoxiao Jiang, Suyi Li, Lingyun Yang +12

Generative image editing using diffusion models has become a prevalent application in today's AI cloud services. In production environments, image editing typically involves a mask…

cs.DC2025

Adaptra: Straggler-Resilient Hybrid-Parallel Training with Pipeline Adaptation

Tianyuan Wu, Lunxi Cao, Hanfeng Lu +8

Training large Deep Neural Network (DNN) models at scale often encounters straggler issues, mostly in communications due to network congestion, RNIC/switch defects, or topological…

cs.DC20242 cited

FALCON: Pinpointing and Mitigating Stragglers for Large-Scale Hybrid-Parallel Training

Tianyuan Wu, Wei Wang, Yinghao Yu +7

Fail-slows, or stragglers, are common but largely unheeded problems in large-scale hybrid-parallel training that spans thousands of GPU servers and runs for weeks to months. Yet, t…

cs.DC2024

SwiftDiffusion: Efficient Diffusion Model Serving with Add-on Modules

Suyi Li, Lingyun Yang, Xiaoxiao Jiang +12

Text-to-image (T2I) generation using diffusion models has become a blockbuster service in today's AI cloud. A production T2I service typically involves a serving workflow where a b…

cs.DC2017

LERC: Coordinated Cache Management for Data-Parallel Systems

Yinghao Yu, Wei Wang, Jun Zhang +1

Memory caches are being aggressively used in today's data-parallel frameworks such as Spark, Tez and Storm. By caching input and intermediate data in memory, compute tasks can witn…