8 citations · 9 across the 2 of their papers we have counts for
4 papers
Students taught by multimodal teachers are superior action recognizers
Gorjan Radevski, Dusan Grujicic, Matthew Blaschko +2
The focal point of egocentric video understanding is modelling hand-object interactions. Standard models -- CNNs, Vision Transformers, etc. -- which receive RGB frames as input per…
Commands 4 Autonomous Vehicles (C4AV) Workshop Summary
Thierry Deruyttere, Simon Vandenhende, Dusan Grujicic +5
The task of visual grounding requires locating the most relevant region or object in an image, given a natural language query. So far, progress on this task was mostly measured on…
A Baseline for the Commands For Autonomous Vehicles Challenge
Simon Vandenhende, Thierry Deruyttere, Dusan Grujicic
The Commands For Autonomous Vehicles (C4AV) challenge requires participants to solve an object referral task in a real-world setting. More specifically, we consider a scenario wher…
Talk2Car: Taking Control of Your Self-Driving Car
Thierry Deruyttere, Simon Vandenhende, Dusan Grujicic +2
A long-term goal of artificial intelligence is to have an agent execute commands communicated through natural language. In many cases the commands are grounded in a visual environm…