2 papers
cs.CV2026
Feature Evolution and Migration during Vision Transformer Training
Joonas Järve, Halil Ibrahim Aysel, Tarun Khajuria +1
We present a novel view on feature evolution in Vision Transformers (ViTs) by visualizing the training process over two dimensions -- network depth (layer) and training time (epoch…
cs.CV2024
Interpreting the structure of multi-object representations in vision encoders
Tarun Khajuria, Braian Olmiro Dias, Marharyta Domnich +1
In this work, we interpret the representations of multi-object scenes in vision encoders through the lens of structured representations. Structured representations allow modeling o…