2 papers
cs.DC2025
Improving training time and GPU utilization in geo-distributed language model training
Palak, Tella Rajashekhar Reddy, Bhaskar Kataria +4
The widespread adoption of language models (LMs) has caused a huge surge in demand for GPUs. Training large LMs requires tens of thousands of GPUs and housing them in the same data…
cs.DC2025
BeLLMan: Controlling LLM Congestion
Tella Rajashekhar Reddy, Atharva Deshmukh, Karan Tandon +3
Large language model (LLM) applications are blindfolded to the infrastructure underneath and generate tokens autoregressively, indifferent to the system load, thus risking inferenc…