papers

Publications (6)

cs.LG2022

Limitations of Neural Collapse for Understanding Generalization in Deep Learning

Like Hui, Mikhail Belkin, Preetum Nakkiran

The recent work of Papyan, Han, & Donoho (2020) presented an intriguing "Neural Collapse" phenomenon, showing a structural property of interpolating classifiers in the late stage o…

cs.LG2021

Evaluation of Neural Architectures Trained with Square Loss vs Cross-Entropy in Classification Tasks

Like Hui, Mikhail Belkin

Modern neural architectures for classification tasks are trained using the cross-entropy loss, which is widely believed to be empirically superior to the square loss. In this work…

cs.LG2023

Cut your Losses with Squentropy

Like Hui, Mikhail Belkin, Stephen Wright

Nearly all practical neural models for classification are trained using cross-entropy loss. Yet this ubiquitous choice is supported by little theoretical or empirical evidence. Rec…

cs.CL2026

Consilience for Verifier-Free Test-Time Scaling

Lecheng Kong, Like Hui, Haitao Mao +1

Test-time scaling often uses an external verifier, such as compilers and test cases in coding or trained value functions in robotics applications, to obtain high-quality rollouts.…

cs.LG2025

Better NTK Conditioning: A Free Lunch from (ReLU) Nonlinear Activation in Wide Neural Networks

Chaoyue Liu, Han Bi, Like Hui +1

Nonlinear activation functions are widely recognized for enhancing the expressivity of neural networks, which is the primary reason for their widespread implementation. In this wor…

cs.LG2018

Kernel Machines Beat Deep Neural Networks on Mask-based Single-channel Speech Enhancement

Like Hui, Siyuan Ma, Mikhail Belkin

We apply a fast kernel method for mask-based single-channel speech enhancement. Specifically, our method solves a kernel regression problem associated to a non-smooth kernel functi…