9 papers
Towards Transparent Checkpointing with AI-driven Code Generation
Hai Duc Nguyen, Tekin Bicer, Kyle Chard +2
Adding reliable checkpoint/restart support to an MPI scientific application is a time-consuming expert effort that requires deep knowledge of both the application and resilience. W…
PTStore (Prefix Tensor Store): Distributed Prefix Caching and Replication for High Throughput Inference Serving
Meghana Maghyastha, Robert Underwood, Randal Burns +1
Inspired by the design of client caching in Content Delivery Networks (CDNs), PTStore distributes and replicates popular tensors that form reusable KV cache prefixes, which are the…
Recency/Frequency Adaptive KV Caching for Large Language Model Serving
Yang Shen, Meghana Madhyastha, Robert Underwood +2
Key-value (KV) caching is a powerful technique for accelerating large language model inference and generation. Inference workloads are large and diverse, which makes them difficult…
Beyond Fixed Budgets: Characterizing the Inelasticity and Limitations of Tree-of-Thought Reasoning Strategies
Atkia Mahila, Avinash Maurya, M. Mustafa Rafique +1
Tree of Thought (ToT) search has become a promising direction for improving the reasoning capabilities of large language models, but deploying these methods in practice raises a qu…
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
Moiz Arif, Avinash Maurya, Sudharshan Vazhkudai +1
The transition from standard generative AI to \emph{reasoning-centric architectures}, exemplified by models capable of extensive Chain-of-Thought~(CoT) processing, marks a fundamen…
Deep Optimizer States: Towards Scalable Training of Transformer Models Using Interleaved Offloading
Avinash Maurya, Jie Ye, M. Mustafa Rafique +2
Transformers and large language models~(LLMs) have seen rapid adoption in all domains. Their sizes have exploded to hundreds of billions of parameters and keep increasing. Under th…