◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Yang You

UC Berkeley

23 papers hereh-index 244.6k citations60 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author6
  • last author15

Across the 21 of 23 papers where every author was matched, so the position is known.

fields
  • cs.LG12
  • cs.DC8
  • cs.CV2
  • cs.IR1
affiliations
  • UC Berkeley
Homepage
same name
  • Yang You — 26 papers, h 15
  • Yang You — 20 papers, h 18
  • Yang You — 20 papers, h 8
  • Yang You — 16 papers, h 13
  • Yang You — 15 papers, h 8
  • Yang You — 11 papers, h 5

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20172023
most citedLarge Batch Training of Convolutional Networks

507 citations · 817 across the 20 of their papers we have counts for

collaborators
Showing 2021 · cs.LGShow all

4 papers · 2 filters

cs.LG2021★ 21 cited

Colossal-AI: A Unified Deep Learning System For Large-Scale Parallel Training

Shenggui Li, Hongxin Liu, Zhengda Bian +5

The success of Transformer models has pushed the deep learning model scale to billions of parameters. Due to the limited memory resource of a single GPU, However, the best practice…

cs.LG2021★ 35 cited

PatrickStar: Parallel Training of Pre-trained Models via Chunk-based Memory Management

Jiarui Fang, Zilin Zhu, Shenggui Li +4

The pre-trained model (PTM) is revolutionizing Artificial Intelligence (AI) technology. However, the hardware requirement of PTM training is prohibitively high, making it a game fo…

cs.LG2021★ 4 cited

Sequence Parallelism: Long Sequence Training from System Perspective

Shenggui Li, Fuzhao Xue, Chaitanya Baranwal +2

Transformer achieves promising results on various tasks. However, self-attention suffers from quadratic memory requirements with respect to the sequence length. Existing work focus…

cs.LG2021

An Efficient 2D Method for Training Super-Large Deep Learning Models

Qifan Xu, Shenggui Li, Chaoyu Gong +1

Huge neural network models have shown unprecedented performance in real-world applications. However, due to memory constraints, model parallelism must be utilized to host large mod…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.