activity
20152024
most citedPyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

13 citations · 13 across the 2 of their papers we have counts for

collaborators

5 papers

cs.CL2024

TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training

Wanchao Liang, Tianyu Liu, Less Wright +10

The development of large language models (LLMs) has been instrumental in advancing state-of-the-art natural language processing applications. Training LLMs with billions of paramet…

cs.DC2023★ 13 cited

PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Yanli Zhao, Andrew Gu, Rohan Varma +15

It is widely acknowledged that large models have the potential to deliver superior performance across a broad range of domains. Despite the remarkable progress made in the field of…

cs.DC2018

Supporting Very Large Models using Automatic Dataflow Graph Partitioning

Minjie Wang, Chien-chin Huang, Jinyang Li

This paper presents Tofu, a system that partitions very large DNN models across multiple GPU devices to reduce per-GPU memory footprint. Tofu is designed to partition a dataflow gr…

cs.DC2018

Unifying Data, Model and Hybrid Parallelism in Deep Learning via Tensor Tiling

Minjie Wang, Chien-chin Huang, Jinyang Li

Deep learning systems have become vital tools across many fields, but the increasing model sizes mean that training must be accelerated to maintain such systems' utility. Current s…

cs.CV2015

Get More With Less: Near Real-Time Image Clustering on Mobile Phones

Jorge Ortiz, Chien-Chin Huang, Supriyo Chakraborty

Machine learning algorithms, in conjunction with user data, hold the promise of revolutionizing the way we interact with our phones, and indeed their widespread adoption in the des…