5 papers
Self-attention vector output similarities reveal how machines pay attention
Tal Halevi, Yarden Tzach, Ronit D. Gross +2
The self-attention mechanism has significantly advanced the field of natural language processing, facilitating the development of advanced language-learning machines. Although its…
Learning Mechanism Underlying NLP Pre-Training and Fine-Tuning
Yarden Tzach, Ronit D. Gross, Ella Koresh +4
Natural language processing (NLP) enables the understanding and generation of meaningful human language, typically using a pre-trained complex architecture on a large dataset to le…
Tiny language models
Ronit D. Gross, Yarden Tzach, Tal Halevi +2
A prominent achievement of natural language processing (NLP) is its ability to understand and generate meaningful human language. This capability relies on complex feedforward tran…
Low-latency vision transformers via large-scale multi-head attention
Ronit D. Gross, Tal Halevi, Ella Koresh +2
The emergence of spontaneous symmetry breaking among a few heads of multi-head attention (MHA) across transformer blocks in classification tasks was recently demonstrated through t…
Unified CNNs and transformers underlying learning mechanism reveals multi-head attention modus vivendi
Ella Koresh, Ronit D. Gross, Yuval Meir +3
Convolutional neural networks (CNNs) evaluate short-range correlations in input images which progress along the layers, whereas vision transformer (ViT) architectures evaluate long…