16 citations · 51 across the 15 of their papers we have counts for
3 papers · 1 filter
TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference
Raja Gond, Nipun Kwatra, Ramachandran Ramjee
Distributed inference of large language models (LLMs) using tensor parallelism can introduce communication overheads of % even over GPUs connected via NVLink, a high-speed GPU…
Singularity: Planet-Scale, Preemptive and Elastic Scheduling of AI Workloads
Dharma Shukla, Muthian Sivathanu, Srinidhi Viswanatha +23
Lowering costs by driving high utilization across deep learning workloads is a crucial lever for cloud providers. We present Singularity, Microsoft's globally distributed schedulin…
Varuna: Scalable, Low-cost Training of Massive Deep Learning Models
Sanjith Athlur, Nitika Saran, Muthian Sivathanu +2
Systems for training massive deep learning models (billions of parameters) today assume and require specialized "hyper-clusters": hundreds or thousands of GPUs wired with specializ…