7 papers
Understanding and Improving Communication Performance in Multi-node LLM Inference
Prajwal Singhania, Siddharth Singh, Lannie Dalton Hough +4
As large language models (LLMs) continue to grow in size, distributed inference has become increasingly important. Model-parallel strategies must now efficiently scale not only acr…
Speculating Experts Accelerates Inference for Mixture-of-Experts
Vivan Madan, Prajwal Singhania, Abhinav Bhatele +2
Mixture-of-Experts (MoE) models have gained popularity as a means of scaling the capacity of large language models (LLMs) while maintaining sparse activations and reduced per-token…
Optimizing Agentic Language Model Inference via Speculative Tool Calls
Daniel Nichols, Prajwal Singhania, Charles Jekel +2
Language models (LMs) are becoming increasingly dependent on external tools. LM-based agentic frameworks frequently interact with their environment via such tools to search files,…
Power Law Guided Dynamic Sifting for Efficient Attention
Nirav Koley, Prajwal Singhania, Abhinav Bhatele
Efficient inference on GPUs using large language models remains challenging due to memory bandwidth limitations, particularly during data transfers between High Bandwidth Memory (H…
Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers
Siddharth Singh, Prajwal Singhania, Aditya Ranjan +9
Training and fine-tuning large language models (LLMs) with hundreds of billions to trillions of parameters requires tens of thousands of GPUs, and a highly scalable software stack.…
Loki: Low-rank Keys for Efficient Sparse Attention
Prajwal Singhania, Siddharth Singh, Shwai He +2
Inference on large language models (LLMs) can be expensive in terms of the compute and memory costs involved, especially when long sequence lengths are used. In particular, the sel…