3 papers
cs.LG2026
Greening AI Inference with Accuracy and Latency-aware User Incentives
Vasilios A. Siris, Adamantia Stamou, George D. Stamoulis +2
The widespread use of AI services has raised concerns for its environmental sustainability, towards which recent studies have identified carbon emissions of AI inference as the maj…
cs.DC2025
Large Language Model Partitioning for Low-Latency Inference at the Edge
Dimitrios Kafetzis, Ramin Khalili, Iordanis Koutsopoulos
Large Language Models (LLMs) based on autoregressive, decoder-only Transformers generate text one token at a time, where a token represents a discrete unit of text. As each newly p…
cs.DC2025
Collaborative Split Federated Learning with Parallel Training and Aggregation
Yiannis Papageorgiou, Yannis Thomas, Alexios Filippakopoulos +2
Federated learning (FL) operates based on model exchanges between the server and the clients, and it suffers from significant client-side computation and communication burden. Spli…