collaborators

5 papers

cs.DC2026

AoiZora: Topology-Aware Auto-Parallel Optimization for Inference of Diffusion Transformers

Kaijian Wang, Yuanyuan Xu, Fanjiang Ye +5

Video diffusion has quickly grown into a key generative serving workload, yet producing each clip demands many denoising iterations over large spatio-temporal latents, which puts l…

cs.DC2026

Characterization-Guided GPU Fault Resilience in NVIDIA MPS

Rixin Liu, Xingqi Cui, Kaijian Wang +4

NVIDIA Multi-Process Service (MPS) enables fine-grained GPU sharing by allowing multiple processes to execute concurrently on the same GPU, making it an important mechanism for imp…

cs.LG2026

ALTO: Adaptive LoRA Tuning and Orchestration for Heterogeneous LoRA Training Workloads

Jingwei Zuo, Xinze Feng, Zien Liu +5

Low-Rank Adaptation (LoRA) is now the dominant method for parameter-efficient fine-tuning of large language models, but achieving a high-quality adapter often requires systematic h…

cs.DC2026

GENSERVE: Efficient Co-Serving of Heterogeneous Diffusion Model Workloads

Fanjiang Ye, Zhangke Li, Xinrui Zhong +10

Diffusion models have emerged as the prevailing approach for text-to-image (T2I) and text-to-video (T2V) generation, yet production platforms must increasingly serve both modalitie…

cs.LG2025

SuperGen: An Efficient Ultra-high-resolution Video Generation System with Sketching and Tiling

Fanjiang Ye, Zepeng Zhao, Yi Mu +11

Diffusion models have recently achieved remarkable success in generative tasks (e.g., image and video generation), and the demand for high-quality content (e.g., 2K/4K videos) is r…