◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Surinder Kumar

16 papers hereh-index 5016.8k citations797 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author2
  • last author13

Across the 16 of 16 papers where every author was matched, so the position is known.

fields
  • cs.LG7
  • cs.CL2
  • cond-mat.mes-hall1
  • cond-mat.mtrl-sci1
  • cs.CV1
  • cs.ET1
same name
  • Surinder Kumar — 1 paper, h 5

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20192023
most citedOn the Convergence of Adam and Beyond

1.6k citations · 1.7k across the 12 of their papers we have counts for

collaborators
Showing 2019Show all

2 papers · 1 filter

math.OC2019

Why are Adaptive Methods Good for Attention Models?

Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit +4

While stochastic gradient descent (SGD) is still the \emph{de facto} algorithm in deep learning, adaptive methods like Clipped SGD/Adam have been observed to outperform SGD across…

cs.LG2019★ 1.6k cited

On the Convergence of Adam and Beyond

Sashank J. Reddi, Satyen Kale, Sanjiv Kumar

Several recently proposed stochastic optimization methods that have been successfully used in training deep networks such as RMSProp, Adam, Adadelta, Nadam are based on using gradi…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.