2 citations · 2 across the 2 of their papers we have counts for
3 papers
cs.LG2026
Emergent Capabilities Arise Randomly from Learning Sparse Attention Patterns
Vatsal Baherwani, Zixi Chen, Shikai Qiu +2
Neural scaling laws for transformer language models predict smooth improvements in pretraining loss with increasing parameters, but downstream capabilities such as in-context learn…
cs.LG2026★ 2 cited
From Entropy to Epiplexity: Rethinking Information for Computationally Bounded Intelligence
Marc Finzi, Shikai Qiu, Yiding Jiang +3
Can we learn more from data than existed in the generating process itself? Can new and useful information be constructed from merely applying deterministic transformations to exist…
cs.LG2025
Out-of-Distribution Detection Methods Answer the Wrong Questions
Yucen Lily Li, Daohan Lu, Polina Kirichenko +4
To detect distribution shifts and improve model safety, many out-of-distribution (OOD) detection methods rely on the predictive uncertainty or features of supervised models trained…