2 papers
cs.NI2026
ML-for-ML
Yutong Zhao, Noga H. Rotman, Gianni Antichi +1
AI training workloads are growing rapidly, making their time, energy, and infrastructure costs increasingly important. In shared cloud clusters, training and fine-tuning jobs compe…
cs.DC2025
AQUA: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU Domains
Abhishek Vijaya Kumar, Gianni Antichi, Rachee Singh
Inference on large-language models (LLMs) is constrained by GPU memory capacity. A sudden increase in the number of inference requests to a cloud-hosted LLM can deplete GPU memory,…