1 citations · 2 across the 12 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Hide & Seek: Transformer Symmetries Obscure Sharpness & Riemannian Geometry Finds It
Marvin F. da Silva, Felix Dangel, Sageev Oore
The concept of sharpness has been successfully applied to traditional architectures like MLPs and CNNs to predict their generalization. For transformers, however, recent work repor…
cs.LG2021★ 1 cited
LogAvgExp Provides a Principled and Performant Global Pooling Operator
Scott C. Lowe, Thomas Trappenberg, Sageev Oore
We seek to improve the pooling operation in neural networks, by applying a more theoretically justified operator. We demonstrate that LogSumExp provides a natural OR operator for l…