◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Noam Shazeer

Google

40 papers hereh-index 40249.1k citations146 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • sole author2
  • first author4
  • middle author19
  • last author5

Across the 30 of 40 papers where every author was matched, so the position is known.

fields
  • cs.CL19
  • cs.LG16
  • cs.CV2
  • cs.DC1
  • cs.NE1
  • eess.IV1
affiliations
  • Google
Homepage
same name
  • Noam Shazeer — 2 papers

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20162026
most citedPaLM: Scaling Language Modeling with Pathways

2.1k citations · 4.1k across the 23 of their papers we have counts for

collaborators
Showing 2021Show all

4 papers · 1 filter

cs.LG2021★ 11 cited

Primer: Searching for Efficient Transformers for Language Modeling

David R. So, Wojciech Mańke, Hanxiao Liu +3

Large Transformer models have been central to recent advances in natural language processing. The training and inference costs of these models, however, have grown rapidly and beco…

cs.DC2021

GSPMD: General and Scalable Parallelization for ML Computation Graphs

Yuanzhong Xu, HyoukJoong Lee, Dehao Chen +13

We present GSPMD, an automatic, compiler-based parallelization system for common machine learning computations. It allows users to write programs in the same way as for a single de…

cs.LG2021

Do Transformer Modifications Transfer Across Implementations and Applications?

Sharan Narang, Hyung Won Chung, Yi Tay +13

The research community has proposed copious modifications to the Transformer architecture since it was introduced over three years ago, relatively few of which have seen widespread…

cs.LG2021

Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

William Fedus, Barret Zoph, Noam Shazeer

In deep learning, models typically reuse the same parameters for all inputs. Mixture of Experts (MoE) defies this and instead selects different parameters for each incoming example…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.