◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

A. Varre

4 papers hereh-index 6267 citations10 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author2
  • middle author2

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.LG4

identity via Semantic Scholar / OpenAlex

activity
20242026
collaborators

4 papers

cs.LG2026

(How) Learning Rates Regulate Catastrophic Overtraining

Mark Rofin, Aditya Varre, Nicolas Flammarion

Supervised fine-tuning (SFT) is a common first stage of LLM post-training, teaching the model to follow instructions and shaping its behavior as a helpful assistant. At the same ti…

cs.LG2026

Gradient Flow Polarizes Softmax Outputs towards Low-Entropy Solutions

Aditya Varre, Mark Rofin, Nicolas Flammarion

Understanding the intricate non-convex training dynamics of softmax-based models is crucial for explaining the empirical success of transformers. In this article, we analyze the gr…

cs.LG2025

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points

Aditya Varre, Gizem Yüce, Nicolas Flammarion

Motivated by empirical observations of prolonged plateaus and stage-wise progression during training, we investigate the loss landscape of transformer models trained on in-context…

cs.LG2024

Why Do We Need Weight Decay in Modern Deep Learning?

Francesco D'Angelo, Maksym Andriushchenko, Aditya Varre +1

Weight decay is a broadly used technique for training state-of-the-art deep networks from image classification to large language models. Despite its widespread usage and being exte…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.