2 papers
cs.DC2025
Large Language Model Partitioning for Low-Latency Inference at the Edge
Dimitrios Kafetzis, Ramin Khalili, Iordanis Koutsopoulos
Large Language Models (LLMs) based on autoregressive, decoder-only Transformers generate text one token at a time, where a token represents a discrete unit of text. As each newly p…
cs.DC2025
Collaborative Split Federated Learning with Parallel Training and Aggregation
Yiannis Papageorgiou, Yannis Thomas, Alexios Filippakopoulos +2
Federated learning (FL) operates based on model exchanges between the server and the clients, and it suffers from significant client-side computation and communication burden. Spli…