◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Jeremy Bernstein

MIT

3 papers hereh-index 142.9k citations24 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author2

Across the 2 of 3 papers where every author was matched, so the position is known.

fields
  • cs.LG3
affiliations
  • MIT
Homepage

identity via Semantic Scholar / OpenAlex

collaborators

3 papers

cs.LG2025

Training Transformers with Enforced Lipschitz Constants

Laker Newhouse, R. Preston Hess, Franz Cesista +3

Neural networks are often highly sensitive to input and weight perturbations. This sensitivity has been linked to pathologies such as vulnerability to adversarial examples, diverge…

cs.LG2024

Modular Duality in Deep Learning

Jeremy Bernstein, Laker Newhouse

An old idea in optimization theory says that since the gradient is a dual vector it may not be subtracted from the weights without first being mapped to the primal space where the…

cs.LG2024

Old Optimizer, New Norm: An Anthology

Jeremy Bernstein, Laker Newhouse

Deep learning optimizers are often motivated through a mix of convex and approximate second-order theory. We select three such methods -- Adam, Shampoo and Prodigy -- and argue tha…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.