3 papers
cs.LG2026
Benign Loss Landscapes Can Coexist with Worst-Case Hardness
Zach Furman, Stephan Wäldchen, Yangda Bei +1
Deep neural networks are expressive enough to contain worst-case targets that can be evaluated in polynomial time but cannot be learned in polynomial time by gradient descent. For…
cs.LG2024
Estimating the Local Learning Coefficient at Scale
Zach Furman, Edmund Lau
The \textit{local learning coefficient} (LLC) is a principled way of quantifying model complexity, originally derived in the context of Bayesian statistics using singular learning…
cs.LG2023
Eliciting Latent Predictions from Transformers with the Tuned Lens
Nora Belrose, Igor Ostrovsky, Lev McKinney +5
We analyze transformers from the perspective of iterative inference, seeking to understand how model predictions are refined layer by layer. To do so, we train an affine probe for…