2 papers
cs.LG2023
Cut your Losses with Squentropy
Like Hui, Mikhail Belkin, Stephen Wright
Nearly all practical neural models for classification are trained using cross-entropy loss. Yet this ubiquitous choice is supported by little theoretical or empirical evidence. Rec…
math.OC2023
A Fully First-Order Method for Stochastic Bilevel Optimization
Jeongyeol Kwon, Dohyun Kwon, Stephen Wright +1
We consider stochastic unconstrained bilevel optimization problems when only the first-order gradient oracles are available. While numerous optimization methods have been proposed…