4 papers
Ampere: Communication-Efficient and High-Accuracy Split Federated Learning
Zihan Zhang, Leon Wong, Blesson Varghese
A Federated Learning (FL) system collaboratively trains neural networks across devices and a server but is limited by significant on-device computation costs. Split Federated Learn…
FedOptima: Optimizing Resource Utilization in Federated Learning
Zihan Zhang, Leon Wong, Blesson Varghese
Federated learning (FL) systems facilitate distributed machine learning across a server and multiple devices. However, FL systems have low resource utilization on servers and devic…
Mosaic: Composite Projection Pruning for Resource-efficient LLMs
Bailey J. Eccles, Leon Wong, Blesson Varghese
Extensive compute and memory requirements limit the deployment of large language models (LLMs) on any hardware. Compression methods, such as pruning, can reduce model size, which i…
Rapid Deployment of DNNs for Edge Computing via Structured Pruning at Initialization
Bailey J. Eccles, Leon Wong, Blesson Varghese
Edge machine learning (ML) enables localized processing of data on devices and is underpinned by deep neural networks (DNNs). However, DNNs cannot be easily run on devices due to t…