2 papers
cs.LG2026
Gradient Boosting within a Single Attention Layer
Saleh Sargolzaei
Transformer attention computes a single softmax-weighted average over values -- a one-pass estimate that cannot correct its own errors. We introduce \emph{gradient-boosted attentio…
cs.CV2024
Improving Out-of-Distribution Data Handling and Corruption Resistance via Modern Hopfield Networks
Saleh Sargolzaei, Luis Rueda
This study explores the potential of Modern Hopfield Networks (MHN) in improving the ability of computer vision models to handle out-of-distribution data. While current computer vi…