most citedDeep Optimizer States: Towards Scalable Training of Transformer Models Using Interleaved Offloading

9 citations · 9 across the 2 of their papers we have counts for

collaborators

8 papers

cs.LG2026

BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models

Zhengyang Wang, Ziyue Liu, Ruijie Zhang +5

The scale of transformer model pre-training is constrained by the increasing computation and communication cost. Low-rank bottleneck architectures offer a promising solution to sig…

cs.LG20269 cited

Deep Optimizer States: Towards Scalable Training of Transformer Models Using Interleaved Offloading

Avinash Maurya, Jie Ye, M. Mustafa Rafique +2

Transformers and large language models~(LLMs) have seen rapid adoption in all domains. Their sizes have exploded to hundreds of billions of parameters and keep increasing. Under th…

cs.DC2026

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers

Avinash Maurya, M. Mustafa Rafique, Franck Cappello +1

The rapid growth of Large Transformer-based models, specifically Large Language Models (LLMs), now scaling to trillions of parameters, has necessitated training across thousands of…

cs.LG2025

CoLA: Compute-Efficient Pre-Training of LLMs via Low-Rank Activation

Ziyue Liu, Ruijie Zhang, Zhengyang Wang +7

The full-size MLPs and the projection layers in attention introduce tremendous model sizes of large language models (LLMs), consuming extensive computational resources in pre-train…

cs.DC2025

MLP-Offload: Multi-Level, Multi-Path Offloading for LLM Pre-training to Break the GPU Memory Wall

Avinash Maurya, M. Mustafa Rafique, Franck Cappello +1

Training LLMs larger than the aggregated memory of multiple GPUs is increasingly necessary due to the faster growth of LLM sizes compared to GPU memory. To this end, multi-tier hos…

cs.AI2025

EAIRA: Establishing a Methodology for Evaluating AI Models as Scientific Research Assistants

Franck Cappello, Sandeep Madireddy, Robert Underwood +23

Recent advancements have positioned AI, and particularly Large Language Models (LLMs), as transformative tools for scientific research, capable of addressing complex tasks that req…