8 citations · 9 across the 6 of their papers we have counts for
4 papers · 1 filter
Eff-GRot: Efficient and Generalizable Rotation Estimation with Transformers
Fanis Mathioulakis, Gorjan Radevski, Tinne Tuytelaars
We introduce Eff-GRot, an approach for efficient and generalizable rotation estimation from RGB images. Given a query image and a set of reference images with known orientations, o…
DAVE: Diagnostic benchmark for Audio Visual Evaluation
Gorjan Radevski, Teodora Popordanoska, Matthew B. Blaschko +1
Audio-visual understanding is a rapidly evolving field that seeks to integrate and interpret information from both auditory and visual modalities. Despite recent advances in multi-…
Students taught by multimodal teachers are superior action recognizers
Gorjan Radevski, Dusan Grujicic, Matthew Blaschko +2
The focal point of egocentric video understanding is modelling hand-object interactions. Standard models -- CNNs, Vision Transformers, etc. -- which receive RGB frames as input per…
Revisiting spatio-temporal layouts for compositional action recognition
Gorjan Radevski, Marie-Francine Moens, Tinne Tuytelaars
Recognizing human actions is fundamentally a spatio-temporal reasoning problem, and should be, at least to some extent, invariant to the appearance of the human and the objects inv…