6 citations · 8 across the 3 of their papers we have counts for
4 papers · 1 filter
Exploring High-Order Self-Similarity for Video Understanding
Manjin Kim, Heeseung Kwon, Karteek Alahari +1
Space-time self-similarity (STSS), which captures visual correspondences across frames, provides an effective way to represent temporal dynamics for video understanding. In this wo…
Learning Correlation Structures for Vision Transformers
Manjin Kim, Paul Hongsuck Seo, Cordelia Schmid +1
We introduce a new attention mechanism, dubbed structural self-attention (StructSA), that leverages rich correlation patterns naturally emerging in key-query interactions of attent…
Relational Self-Attention: What's Missing in Attention for Video Understanding
Manjin Kim, Heeseung Kwon, Chunyu Wang +2
Convolution has been arguably the most important feature transform for modern neural networks, leading to the advance of deep learning. Recent emergence of Transformer networks, wh…
Learning Self-Similarity in Space and Time as Generalized Motion for Video Action Recognition
Heeseung Kwon, Manjin Kim, Suha Kwak +1
Spatio-temporal convolution often fails to learn motion dynamics in videos and thus an effective motion representation is required for video understanding in the wild. In this pape…