◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Li Dong

Microsoft Research

55 papers hereh-index 6331.4k citations118 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author3
  • middle author47

Across the 50 of 55 papers where every author was matched, so the position is known.

fields
  • cs.CL45
  • cs.CV6
  • cs.LG4
affiliations
  • Microsoft Research
Homepage
same name
  • Li Dong — 33 papers, h 11
  • Li Dong — 25 papers, h 18
  • Li Dong — 6 papers, h 5
  • Li Dong — 5 papers, h 10
  • Li Dong — 4 papers, h 3
  • Li Dong — 4 papers, h 3

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20162026
most citedUniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training

225 citations · 791 across the 33 of their papers we have counts for

collaborators
Showing cs.LGShow all

4 papers · 1 filter

cs.LG2022★ 6 cited

TorchScale: Transformers at Scale

Shuming Ma, Hongyu Wang, Shaohan Huang +8

Large Transformers have achieved state-of-the-art performance across many tasks. Most open-source libraries on scaling Transformers focus on improving training or inference with be…

cs.LG2022★ 13 cited

Foundation Transformers

Hongyu Wang, Shuming Ma, Shaohan Huang +12

A big convergence of model architectures across language, vision, speech, and multimodal is emerging. However, under the same name "Transformers", the above areas use different imp…

cs.LG2022

StableMoE: Stable Routing Strategy for Mixture of Experts

Damai Dai, Li Dong, Shuming Ma +4

The Mixture-of-Experts (MoE) technique can scale up the model size of Transformers with an affordable computational overhead. We point out that existing learning-to-route MoE metho…

cs.LG2021★ 2 cited

Memory-Efficient Differentiable Transformer Architecture Search

Yuekai Zhao, Li Dong, Yelong Shen +3

Differentiable architecture search (DARTS) is successfully applied in many vision tasks. However, directly using DARTS for Transformers is memory-intensive, which renders the searc…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.