◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Surinder Kumar

3 papers hereh-index 5016.8k citations797 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author1
  • last author2

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.LG2
  • math.OC1

identity via Semantic Scholar / OpenAlex

most citedOn the Convergence of Adam and Beyond

1.6k citations · 1.6k across the 2 of their papers we have counts for

collaborators

3 papers

cs.LG2020★ 18 cited

Does label smoothing mitigate label noise?

Michal Lukasik, Srinadh Bhojanapalli, Aditya Krishna Menon +1

Label smoothing is commonly used in training deep learning models, wherein one-hot training labels are mixed with uniform label vectors. Empirically, smoothing has been shown to im…

math.OC2019

Why are Adaptive Methods Good for Attention Models?

Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit +4

While stochastic gradient descent (SGD) is still the \emph{de facto} algorithm in deep learning, adaptive methods like Clipped SGD/Adam have been observed to outperform SGD across…

cs.LG2019★ 1.6k cited

On the Convergence of Adam and Beyond

Sashank J. Reddi, Satyen Kale, Sanjiv Kumar

Several recently proposed stochastic optimization methods that have been successfully used in training deep networks such as RMSProp, Adam, Adadelta, Nadam are based on using gradi…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.