1 paper
Ryan Grainger, Thomas Paniagua, Xi Song +3
Vision Transformers (ViTs) are built on the assumption of treating image patches as ``visual tokens" and learn patch-to-patch attention. The patch embedding based tokenizer has a s…