Publications (6)
Limitations of Neural Collapse for Understanding Generalization in Deep Learning
Like Hui, Mikhail Belkin, Preetum Nakkiran
The recent work of Papyan, Han, & Donoho (2020) presented an intriguing "Neural Collapse" phenomenon, showing a structural property of interpolating classifiers in the late stage o…
Evaluation of Neural Architectures Trained with Square Loss vs Cross-Entropy in Classification Tasks
Like Hui, Mikhail Belkin
Modern neural architectures for classification tasks are trained using the cross-entropy loss, which is widely believed to be empirically superior to the square loss. In this work…
Cut your Losses with Squentropy
Like Hui, Mikhail Belkin, Stephen Wright
Nearly all practical neural models for classification are trained using cross-entropy loss. Yet this ubiquitous choice is supported by little theoretical or empirical evidence. Rec…
Consilience for Verifier-Free Test-Time Scaling
Lecheng Kong, Like Hui, Haitao Mao +1
Test-time scaling often uses an external verifier, such as compilers and test cases in coding or trained value functions in robotics applications, to obtain high-quality rollouts.…
Better NTK Conditioning: A Free Lunch from (ReLU) Nonlinear Activation in Wide Neural Networks
Chaoyue Liu, Han Bi, Like Hui +1
Nonlinear activation functions are widely recognized for enhancing the expressivity of neural networks, which is the primary reason for their widespread implementation. In this wor…
Kernel Machines Beat Deep Neural Networks on Mask-based Single-channel Speech Enhancement
Like Hui, Siyuan Ma, Mikhail Belkin
We apply a fast kernel method for mask-based single-channel speech enhancement. Specifically, our method solves a kernel regression problem associated to a non-smooth kernel functi…