65 citations · 73 across the 10 of their papers we have counts for
10 papers
Even Sparser Graph Transformers
Hamed Shirzad, Honghao Lin, Balaji Venkatachalam +3
Graph Transformers excel in long-range dependency modeling, but generally require quadratic memory complexity in the number of nodes in an input graph, and hence have trouble scali…
A Theory for Compressibility of Graph Transformers for Transductive Learning
Hamed Shirzad, Honghao Lin, Ameya Velingker +3
Transductive tasks on graphs differ fundamentally from typical supervised machine learning tasks, as the independent and identically distributed (i.i.d.) assumption does not hold a…
Generalized Coverage for More Robust Low-Budget Active Learning
Wonho Bae, Junhyug Noh, Danica J. Sutherland
The ProbCover method of Yehuda et al. is a well-motivated algorithm for active learning in low-budget regimes, which attempts to "cover" the data distribution with balls of a given…
Why Do You Grok? A Theoretical Analysis of Grokking Modular Addition
Mohamad Amin Mohamadi, Zhiyuan Li, Lei Wu +1
We present a theoretical explanation of the ``grokking'' phenomenon, where a model generalizes long after overfitting,for the originally-studied problem of modular addition. First,…
AdaFlood: Adaptive Flood Regularization
Wonho Bae, Yi Ren, Mohamad Osama Ahmed +3
Although neural networks are conventionally optimized towards zero training loss, it has been recently learned that targeting a non-zero training loss threshold, referred to as a f…
Improving Compositional Generalization Using Iterated Learning and Simplicial Embeddings
Yi Ren, Samuel Lavoie, Mikhail Galkin +2
Compositional generalization, the ability of an agent to generalize to unseen combinations of latent factors, is easy for humans but hard for deep neural networks. A line of resear…