3 citations · 5 across the 2 of their papers we have counts for
4 papers
Convexifying Transformers: Improving optimization and understanding of transformer networks
Tolga Ergen, Behnam Neyshabur, Harsh Mehta
Understanding the fundamental mechanism behind the success of transformer networks is still an open problem in the deep learning literature. Although their remarkable performance h…
Unraveling Attention via Convex Duality: Analysis and Interpretations of Vision Transformers
Arda Sahiner, Tolga Ergen, Batu Ozturkler +3
Vision transformers using self-attention or its proposed alternatives have demonstrated promising results in many image related tasks. However, the underpinning inductive bias of a…
Implicit Convex Regularizers of CNN Architectures: Convex Optimization of Two- and Three-Layer Networks in Polynomial Time
Tolga Ergen, Mert Pilanci
We study training of Convolutional Neural Networks (CNNs) with ReLU activations and introduce exact convex optimization formulations with a polynomial complexity with respect to th…
Neural Networks are Convex Regularizers: Exact Polynomial-time Convex Optimization Formulations for Two-layer Networks
Mert Pilanci, Tolga Ergen
We develop exact representations of training two-layer neural networks with rectified linear units (ReLUs) in terms of a single convex program with number of variables polynomial i…