3 papers
cs.DC2025
Spatio-Temporal Parallelism for Diffusion Model Inference on Heterogeneous Multi-GPU Systems
Han Liang, Jiahui Zhou, Zicheng Zhou +2
The widespread adoption of diffusion models for image generation necessitates efficient parallel inference to manage their substantial computational overhead. However, current para…
cs.LG2025
SmallThinker: A Family of Efficient Large Language Models Natively Trained for Local Deployment
Yixin Song, Zhenliang Xue, Dongliang Wei +11
While frontier large language models (LLMs) continue to push capability boundaries, their deployment remains confined to GPU-powered cloud infrastructure. We challenge this paradig…
cs.LG2024
FedReMa: Improving Personalized Federated Learning via Leveraging the Most Relevant Clients
Han Liang, Ziwei Zhan, Weijie Liu +3
Federated Learning (FL) is a distributed machine learning paradigm that achieves a globally robust model through decentralized computation and periodic model synthesis, primarily f…