3 papers
cs.CV2025
DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition
Hayat Ullah, Muhammad Ali Shafique, Abbas Khan +1
The landscape of video recognition has evolved significantly, shifting from traditional Convolutional Neural Networks (CNNs) to Transformer-based architectures for improved accurac…
cs.CV2025
OD-VIRAT: A Large-Scale Benchmark for Object Detection in Realistic Surveillance Environments
Hayat Ullah, Abbas Khan, Arslan Munir +1
Realistic human surveillance datasets are crucial for training and evaluating computer vision models under real-world conditions, facilitating the development of robust algorithms…
cs.CV2025
Hierarchical Multi-Stage Transformer Architecture for Context-Aware Temporal Action Localization
Hayat Ullah, Arslan Munir, Oliver Nina
Inspired by the recent success of transformers and multi-stage architectures in video recognition and object detection domains. We thoroughly explore the rich spatio-temporal prope…