collaborators

7 papers

cs.LG2026

PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint

Haotian Xie, Junlin Chen, Mingkai Zheng +2

State-of-the-art large language model (LLM) training takes tens of thousands of graphics processing units (GPUs) for months and encounters failures across the software and hardware…

cs.LG2026

FediLoRA: Practical Federated Fine-Tuning of Foundation Models Under Missing-Modality Constraints

Lishan Yang, Wei Emma Zhang, Nam Kha Nguygen +4

Federated Learning with LoRA fine-tuning offers an efficient and privacy-aware solution for institutions to collaboratively leverage their large datasets to train VLLMs. However, p…

cs.CR2025

PrisonBreak: Jailbreaking Large Language Models with at Most Twenty-Five Targeted Bit-flips

Zachary Coalson, Jeonghyun Woo, Chris S. Lin +8

We study a new vulnerability in commercial-scale safety-aligned large language models (LLMs): their refusal to generate harmful responses can be broken by flipping only a few bits…

cs.DC2025

Understanding the Landscape of Ampere GPU Memory Errors

Zhu Zhu, Yu Sun, Dhatri Parakal +9

Graphics Processing Units (GPUs) have become a de facto solution for accelerating high-performance computing (HPC) applications. Understanding their memory error behavior is an ess…

cs.LG2025

MMiC: Mitigating Modality Incompleteness in Clustered Federated Learning

Lishan Yang, Wei Emma Zhang, Quan Z. Sheng +3

In the era of big data, data mining has become indispensable for uncovering hidden patterns and insights from vast and complex datasets. The integration of multimodal data sources…

cs.LG2025

FedDPG: An Adaptive Yet Efficient Prompt-tuning Approach in Federated Learning Settings

Ali Shakeri, Wei Emma Zhang, Amin Beheshti +3

Pre-trained Language Models (PLMs) have demonstrated impressive performance in various NLP tasks. However, traditional fine-tuning methods for leveraging PLMs for downstream tasks…