1 paper · 1 filter
Gorjan Radevski, Dusan Grujicic, Marie-Francine Moens +2
The focal point of egocentric video understanding is modelling hand-object interactions. Standard models, e.g. CNNs or Vision Transformers, which receive RGB frames as input perfor…