5 papers · 1 filter
Reinforcement Learning meets Masked Video Modeling : Trajectory-Guided Adaptive Token Selection
Ayush K. Rai, Kyle Min, Tarun Krishna +3
Masked video modeling~(MVM) has emerged as a highly effective pre-training strategy for visual foundation models, whereby the model reconstructs masked spatiotemporal tokens using…
Resilience of Vision Transformers for Domain Generalisation in the Presence of Out-of-Distribution Noisy Images
Hamza Riaz, Alan F. Smeaton
Modern AI models excel in controlled settings but often fail in real-world scenarios where data distributions shift unpredictably - a challenge known as domain generalisation (DG).…
The Effects of Grouped Structural Global Pruning of Vision Transformers on Domain Generalisation
Hamza Riaz, Alan F. Smeaton
With the growing sizes of AI models like large language models (LLMs) and vision transformers, deploying them on devices with limited computational resources is a significant chall…
Efficient Object-centric Representation Learning with Pre-trained Geometric Prior
Phúc H. Le Khac, Graham Healy, Alan F. Smeaton
This paper addresses key challenges in object-centric representation learning of video. While existing approaches struggle with complex scenes, we propose a novel weakly-supervised…
Generative Outpainting To Enhance the Memorability of Short-Form Videos
Alan Byju, Aman Sudhindra Ladwa, Lorin Sweeney +1
With the expanding use of the short-form video format in advertising, social media, entertainment, education and more, there is a need for such media to both captivate and be remem…