activity
20242026
most citedTowards understanding how attention mechanism works in deep learning

5 citations · 5 across the 10 of their papers we have counts for

collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG2026

Muon Learns More Robust and Transferable Features than Adam

Tianyu Ruan, Fengzhuo Zhang, Shuche Wang +1

Muon has recently emerged as a state-of-the-art optimizer for pretraining Large Language Models (LLMs) and vision classifiers. Despite its efficiency advantage over Adam and SGD, t…

cs.LG2026

Beyond Neural Collapse: Task-Intrinsic Geometry Governs Neural Representations in Modular Arithmetic

Hu Tan, Kuo Gai, Shihua Zhang

While neural collapse (NC) predicts that a -class-balanced classifier should organize terminal representations as a -dimensional simplex equiangular tight frame (ETF), mo…

cs.LG2026

Deciphering Two Training Clocks in Grokking via Deep Linear Network Theory with Conditional ReLU Reduction

Hu Tan, Kuo Gai, Shihua Zhang

Grokking suggests that fitting the training data and learning a simple underlying rule may occur on different time scales. We formalize this phenomenon by separating the fast decay…

cs.LG2025

Feature Dynamics as Implicit Data Augmentation: A Depth-Decomposed View on Deep Neural Network Generalization

Tianyu Ruan, Kuo Gai, Shihua Zhang

Why do deep networks generalize well? In contrast to classical generalization theory, we approach this fundamental question by examining not only inputs and outputs, but the evolut…

cs.LG2024★ 5 cited

Towards understanding how attention mechanism works in deep learning

Tianyu Ruan, Shihua Zhang

Attention mechanism has been extensively integrated within mainstream neural network architectures, such as Transformers and graph attention networks. Yet, its underlying working p…

cs.LG2024

OTAD: An Optimal Transport-Induced Robust Model for Agnostic Adversarial Attack

Kuo Gai, Sicong Wang, Shihua Zhang

Deep neural networks (DNNs) are vulnerable to small adversarial perturbations of the inputs, posing a significant challenge to their reliability and robustness. Empirical methods s…