activity
20172026
most citedWhy Normalizing Flows Fail to Detect Out-of-Distribution Data

32 citations · 128 across the 10 of their papers we have counts for

collaborators

16 papers

cs.LG2026

Emergent Capabilities Arise Randomly from Learning Sparse Attention Patterns

Vatsal Baherwani, Zixi Chen, Shikai Qiu +2

Neural scaling laws for transformer language models predict smooth improvements in pretraining loss with increasing parameters, but downstream capabilities such as in-context learn…

cs.LG2026

From Entropy to Epiplexity: Rethinking Information for Computationally Bounded Intelligence

Marc Finzi, Shikai Qiu, Yiding Jiang +3

Can we learn more from data than existed in the generating process itself? Can new and useful information be constructed from merely applying deterministic transformations to exist…

cs.LG2025

Out-of-Distribution Detection Methods Answer the Wrong Questions

Yucen Lily Li, Daohan Lu, Polina Kirichenko +4

To detect distribution shifts and improve model safety, many out-of-distribution (OOD) detection methods rely on the predictive uncertainty or features of supervised models trained…

cs.LG202223 cited

On Feature Learning in the Presence of Spurious Correlations

Pavel Izmailov, Polina Kirichenko, Nate Gruver +1

Deep classifiers are known to rely on spurious features $\unicode{x2013}$ patterns which are correlated with the target on the training data but not inherently relevant to the lear…

cs.LG20229 cited

On Uncertainty, Tempering, and Data Augmentation in Bayesian Classification

Sanyam Kapoor, Wesley J. Maddox, Pavel Izmailov +1

Aleatoric uncertainty captures the inherent randomness of the data, such as measurement noise. In Bayesian regression, we often use a Gaussian observation model, where we control t…

cs.LG2021

What Are Bayesian Neural Network Posteriors Really Like?

Pavel Izmailov, Sharad Vikram, Matthew D. Hoffman +1

The posterior over Bayesian neural network (BNN) parameters is extremely high-dimensional and non-convex. For computational reasons, researchers approximate this posterior using in…