2 citations · 3 across the 2 of their papers we have counts for
2 papers
cs.LG2023★ 1 cited
ROAM: memory-efficient large DNN training via optimized operator ordering and memory layout
Huiyao Shu, Ang Wang, Ziji Shi +3
As deep learning models continue to increase in size, the memory requirements for training have surged. While high-level techniques like offloading, recomputation, and compression…
cs.DC2023★ 2 cited
Auto-Parallelizing Large Models with Rhino: A Systematic Approach on Production AI Platform
Shiwei Zhang, Lansong Diao, Siyu Wang +7
We present Rhino, a system for accelerating tensor programs with automatic parallelization on AI platform for real production environment. It transforms a tensor program written fo…