1 citations · 2 across the 2 of their papers we have counts for
4 papers · 1 filter
Snakes and Ladders: Two Steps Up for VideoMamba
Hui Lu, Albert Ali Salah, Ronald Poppe
Video understanding requires the extraction of rich spatio-temporal representations, which transformer models achieve through self-attention. Unfortunately, self-attention poses a…
Enhancing Video Transformers for Action Understanding with VLM-aided Training
Hui Lu, Hu Jian, Ronald Poppe +1
Owing to their ability to extract relevant spatio-temporal video embeddings, Vision Transformers (ViTs) are currently the best performing models in video action understanding. Howe…
TCNet: Continuous Sign Language Recognition from Trajectories and Correlated Regions
Hui Lu, Albert Ali Salah, Ronald Poppe
A key challenge in continuous sign language recognition (CSLR) is to efficiently capture long-range spatial interactions over time from the video input. To address this challenge,…
Compensation Sampling for Improved Convergence in Diffusion Models
Hui Lu, Albert ali Salah, Ronald Poppe
Diffusion models achieve remarkable quality in image generation, but at a cost. Iterative denoising requires many time steps to produce high fidelity images. We argue that the deno…