◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Junyang Lin

28 papers hereh-index 4017.6k citations62 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author7
  • middle author16
  • last author3

Across the 26 of 28 papers where every author was matched, so the position is known.

fields
  • cs.CL21
  • cs.LG4
  • cs.CV3
same name
  • Junyang Lin — 10 papers
  • Junyang Lin — 3 papers
  • Junyang Lin — 1 paper
  • Junyang Lin — 1 paper
  • Junyang Lin — 1 paper

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20182022
most citedUnderstanding and Improving Layer Normalization

178 citations · 343 across the 10 of their papers we have counts for

collaborators
Showing cs.LGShow all

4 papers · 1 filter

cs.LG2022★ 22 cited

Modality Competition: What Makes Joint Training of Multi-modal Network Fail in Deep Learning? (Provably)

Yu Huang, Junyang Lin, Chang Zhou +2

Despite the remarkable success of deep multi-modal learning in practice, it has not been well-explained in theory. Recently, it has been observed that the best uni-modal network ou…

cs.LG2021★ 18 cited

M6-10T: A Sharing-Delinking Paradigm for Efficient Multi-Trillion Parameter Pretraining

Junyang Lin, An Yang, Jinze Bai +9

Recent expeditious developments in deep learning algorithms, distributed training, and even hardware design for large models have enabled training extreme-scale models, say GPT-3 a…

cs.LG2021

M6-T: Exploring Sparse Expert Models and Beyond

An Yang, Junyang Lin, Rui Men +12

Mixture-of-Experts (MoE) models can achieve promising results with outrageous large amount of parameters but constant computation cost, and thus it has become a trend in model scal…

cs.LG2019★ 178 cited

Understanding and Improving Layer Normalization

Jingjing Xu, Xu Sun, Zhiyuan Zhang +2

Layer normalization (LayerNorm) is a technique to normalize the distributions of intermediate layers. It enables smoother gradients, faster training, and better generalization accu…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.