3 papers
cs.LG2026
SNIP: An Adaptive Mixed Precision Framework for Subbyte Large Language Model Training
Yunjie Pan, Yongyi Yang, Hanmei Yang +1
Training large language models (LLMs) efficiently while preserving model quality poses significant challenges, particularly with subbyte precision supported by state-of-the-art GPU…
cs.LG2025
An Equivariance Toolbox for Learning Dynamics
Yongyi Yang, Liu Ziyin
Many theoretical results in deep learning can be traced to symmetry or equivariance of neural networks under parameter transformations. However, existing analyses are typically pro…
cs.LG2025
Topological Invariance and Breakdown in Learning
Yongyi Yang, Tomaso Poggio, Isaac Chuang +1
We prove that for a broad class of permutation-equivariant learning rules (including SGD, Adam, and others), the training process induces a bi-Lipschitz mapping between neurons and…