activity
20242026
collaborators

6 papers

cs.LG2026

Data-Constrained Language Model Pretraining: Improved Regularization and Scaling Laws

Zhiwei Xu, Shihao Wu, Hanseul Cho +2

Classical scaling laws for language model pretraining balance model size against training dataset size under a fixed compute budget, assuming abundant data and a single pass over t…

cs.LG2026

Characterizing Pattern Matching and Its Limits on Compositional Task Structures

Hoyeon Chang, Jinho Park, Hanseul Cho +7

Despite impressive capabilities, LLMs' successes often rely on pattern-matching behaviors, yet these are also linked to OOD generalization failures in compositional tasks. However,…

cs.LG2025

Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification

Hyunji Jung, Hanseul Cho, Chulhee Yun

We study continual learning on multiple linear classification tasks by sequentially running gradient descent (GD) for a fixed budget of iterations per task. When all tasks are join…

cs.LG2025

Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count

Hanseul Cho, Jaeyoung Cha, Srinadh Bhojanapalli +1

Transformers often struggle with length generalization, meaning they fail to generalize to sequences longer than those encountered during training. While arithmetic tasks are commo…

cs.LG2024

DASH: Warm-Starting Neural Network Training in Stationary Settings without Loss of Plasticity

Baekrok Shin, Junsoo Oh, Hanseul Cho +1

Warm-starting neural network training by initializing networks with previously learned weights is appealing, as practical neural networks are often deployed under a continuous infl…

cs.LG2024

Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure

Hanseul Cho, Jaeyoung Cha, Pranjal Awasthi +3

Even for simple arithmetic tasks like integer addition, it is challenging for Transformers to generalize to longer sequences than those encountered during training. To tackle this…