5 citations · 5 across the 10 of their papers we have counts for
7 papers · 1 filter
Muon Learns More Robust and Transferable Features than Adam
Tianyu Ruan, Fengzhuo Zhang, Shuche Wang +1
Muon has recently emerged as a state-of-the-art optimizer for pretraining Large Language Models (LLMs) and vision classifiers. Despite its efficiency advantage over Adam and SGD, t…
Beyond Neural Collapse: Task-Intrinsic Geometry Governs Neural Representations in Modular Arithmetic
Hu Tan, Kuo Gai, Shihua Zhang
While neural collapse (NC) predicts that a -class-balanced classifier should organize terminal representations as a -dimensional simplex equiangular tight frame (ETF), mo…
Deciphering Two Training Clocks in Grokking via Deep Linear Network Theory with Conditional ReLU Reduction
Hu Tan, Kuo Gai, Shihua Zhang
Grokking suggests that fitting the training data and learning a simple underlying rule may occur on different time scales. We formalize this phenomenon by separating the fast decay…
Feature Dynamics as Implicit Data Augmentation: A Depth-Decomposed View on Deep Neural Network Generalization
Tianyu Ruan, Kuo Gai, Shihua Zhang
Why do deep networks generalize well? In contrast to classical generalization theory, we approach this fundamental question by examining not only inputs and outputs, but the evolut…
Towards understanding how attention mechanism works in deep learning
Tianyu Ruan, Shihua Zhang
Attention mechanism has been extensively integrated within mainstream neural network architectures, such as Transformers and graph attention networks. Yet, its underlying working p…
OTAD: An Optimal Transport-Induced Robust Model for Agnostic Adversarial Attack
Kuo Gai, Sicong Wang, Shihua Zhang
Deep neural networks (DNNs) are vulnerable to small adversarial perturbations of the inputs, posing a significant challenge to their reliability and robustness. Empirical methods s…