1 citations · 1 across the 3 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
Low-latency vision transformers via large-scale multi-head attention
Ronit D. Gross, Tal Halevi, Ella Koresh +2
The emergence of spontaneous symmetry breaking among a few heads of multi-head attention (MHA) across transformer blocks in classification tasks was recently demonstrated through t…
cs.CV2023★ 1 cited
The mechanism underlying successful deep learning
Yarden Tzach, Yuval Meir, Ofek Tevet +4
Deep architectures consist of tens or hundreds of convolutional layers (CLs) that terminate with a few fully connected (FC) layers and an output layer representing the possible lab…