32 citations · 128 across the 10 of their papers we have counts for
16 papers
Emergent Capabilities Arise Randomly from Learning Sparse Attention Patterns
Vatsal Baherwani, Zixi Chen, Shikai Qiu +2
Neural scaling laws for transformer language models predict smooth improvements in pretraining loss with increasing parameters, but downstream capabilities such as in-context learn…
From Entropy to Epiplexity: Rethinking Information for Computationally Bounded Intelligence
Marc Finzi, Shikai Qiu, Yiding Jiang +3
Can we learn more from data than existed in the generating process itself? Can new and useful information be constructed from merely applying deterministic transformations to exist…
Out-of-Distribution Detection Methods Answer the Wrong Questions
Yucen Lily Li, Daohan Lu, Polina Kirichenko +4
To detect distribution shifts and improve model safety, many out-of-distribution (OOD) detection methods rely on the predictive uncertainty or features of supervised models trained…
On Feature Learning in the Presence of Spurious Correlations
Pavel Izmailov, Polina Kirichenko, Nate Gruver +1
Deep classifiers are known to rely on spurious features $\unicode{x2013}$ patterns which are correlated with the target on the training data but not inherently relevant to the lear…
On Uncertainty, Tempering, and Data Augmentation in Bayesian Classification
Sanyam Kapoor, Wesley J. Maddox, Pavel Izmailov +1
Aleatoric uncertainty captures the inherent randomness of the data, such as measurement noise. In Bayesian regression, we often use a Gaussian observation model, where we control t…
What Are Bayesian Neural Network Posteriors Really Like?
Pavel Izmailov, Sharad Vikram, Matthew D. Hoffman +1
The posterior over Bayesian neural network (BNN) parameters is extremely high-dimensional and non-convex. For computational reasons, researchers approximate this posterior using in…