8 citations · 22 across the 18 of their papers we have counts for
21 papers · 1 filter
Nearly Tight Rademacher Bounds for Sparsely Activated Neural Networks
Xiaoyu Li, Zhizhou Sha, Jiaojiao Jiang +2
An input may activate few hidden units even when different inputs collectively use an entire network. We study the statistical complexity of this input-dependent sparsity in the on…
Feature Superposition in Neural Networks: From Theory to Practice
Dai Shi, Xiaoyu Li, Andi Han +1
Superposition refers to neural networks representing more features than they have dimensions. It offers a possible explanation for polysemantic neurons and motivates methods for re…
Optimistic Rates for Multiclass PAC Learning
Xiaoyu Li, Andi Han, Jiaojiao Jiang +1
Worst-case multiclass bounds do not become smaller when the best classifier is already nearly correct: what is missing is an optimistic rate, a guarantee whose fluctuation scales w…
Flood and Harvest: The Provable Necessity of Trivia for Generating Valuable Mathematics via the Lens of Language Generation in the Limit
Xiaoyu Li, Andi Han, Dai Shi +3
AI systems coupled to proof assistants now generate formal mathematics at scale, and the gap between what a checker can verify and what a mathematician would value has become the b…
Learning Manifold and Itô Dynamics with Branched Neural Rough Differential Equations
Luke Thompson, Dai Shi, Lequan Lin +2
Neural rough differential equations (NRDEs) stay accurate under irregular sampling while taking far fewer integration steps than standard neural differential equations, summarising…
SGNN: Efficient Global Mixing and Local Message Passing for Long-Range Graph Learning
Dai Shi, Luke Thompson, Linhan Luo +4
Message-passing neural networks (MPNNs) often suffer from an information bottleneck when capturing long-range dependencies, leading to the oversquashing (OSQ) phenomenon. Alongside…