3 papers
cs.CV2024
Snakes and Ladders: Two Steps Up for VideoMamba
Hui Lu, Albert Ali Salah, Ronald Poppe
Video understanding requires the extraction of rich spatio-temporal representations, which transformer models achieve through self-attention. Unfortunately, self-attention poses a…
cs.CV2024
Enhancing Video Transformers for Action Understanding with VLM-aided Training
Hui Lu, Hu Jian, Ronald Poppe +1
Owing to their ability to extract relevant spatio-temporal video embeddings, Vision Transformers (ViTs) are currently the best performing models in video action understanding. Howe…
cs.CV2024
TCNet: Continuous Sign Language Recognition from Trajectories and Correlated Regions
Hui Lu, Albert Ali Salah, Ronald Poppe
A key challenge in continuous sign language recognition (CSLR) is to efficiently capture long-range spatial interactions over time from the video input. To address this challenge,…