1 paper
Vivek Trivedy, Amani Almalki, Longin Jan Latecki
We propose an adaptation to the training of Vision Transformers (ViTs) that allows for an explicit modeling of objects during the attention computation. This is achieved by adding…