40 citations · 60 across the 13 of their papers we have counts for
Showing 2021 · cs.CVShow all
3 papers · 2 filters
cs.CV2021
Leveraging Local Temporal Information for Multimodal Scene Classification
Saurabh Sahu, Palash Goyal
Robust video scene classification models should capture the spatial (pixel-wise) and temporal (frame-wise) characteristics of a video effectively. Transformer models with self-atte…
cs.CV2021
Can't Fool Me: Adversarially Robust Transformer for Video Understanding
Divya Choudhary, Palash Goyal, Saurabh Sahu
Deep neural networks have been shown to perform poorly on adversarial examples. To address this, several techniques have been proposed to increase robustness of a model for image c…
cs.CV2021★ 1 cited
Enhancing Transformer for Video Understanding Using Gated Multi-Level Attention and Temporal Adversarial Training
Saurabh Sahu, Palash Goyal
The introduction of Transformer model has led to tremendous advancements in sequence modeling, especially in text domain. However, the use of attention-based models for video under…