2 papers
cs.LG2026
DART-FL: Burst-Aware Multitask Federated Learning under Dynamic Inference Demand at the Edge
Yiming Xie, Pinrui Yu, Geng Yuan +2
Edge intelligence systems increasingly require model training and online inference to coexist on resource-constrained devices, while inference demand can vary substantially across…
cs.LG2024
Pruning Foundation Models for High Accuracy without Retraining
Pu Zhao, Fei Sun, Xuan Shen +4
Despite the superior performance, it is challenging to deploy foundation models or large language models (LLMs) due to their massive parameters and computations. While pruning is a…