1 citations · 1 across the 3 of their papers we have counts for
3 papers
Composition of Experts: A Modular Compound AI System Leveraging Large Language Models
Swayambhoo Jain, Ravi Raju, Bo Li +8
Large Language Models (LLMs) have achieved remarkable advancements, but their monolithic nature presents challenges in terms of scalability, cost, and customization. This paper int…
Kernel Looping: Eliminating Synchronization Boundaries for Peak Inference Performance
David Koeplinger, Darshan Gandhi, Pushkar Nandkar +9
Token generation speed is critical to power the next wave of AI inference applications. GPUs significantly underperform during token generation due to synchronization overheads at…
Training Large Language Models Efficiently with Sparsity and Dataflow
Venkat Srinivasan, Darshan Gandhi, Urmish Thakker +1
Large foundation language models have shown their versatility in being able to be adapted to perform a wide variety of downstream tasks, such as text generation, sentiment analysis…