activity
20182021
most citedEnhancing Transformer for Video Understanding Using Gated Multi-Level Attention and Temporal Adversarial Training

1 citations · 1 across the 3 of their papers we have counts for

collaborators

9 papers

cs.CV2021

Leveraging Local Temporal Information for Multimodal Scene Classification

Saurabh Sahu, Palash Goyal

Robust video scene classification models should capture the spatial (pixel-wise) and temporal (frame-wise) characteristics of a video effectively. Transformer models with self-atte…

cs.CV2021

Can't Fool Me: Adversarially Robust Transformer for Video Understanding

Divya Choudhary, Palash Goyal, Saurabh Sahu

Deep neural networks have been shown to perform poorly on adversarial examples. To address this, several techniques have been proposed to increase robustness of a model for image c…

cs.CV20211 cited

Enhancing Transformer for Video Understanding Using Gated Multi-Level Attention and Temporal Adversarial Training

Saurabh Sahu, Palash Goyal

The introduction of Transformer model has led to tremendous advancements in sequence modeling, especially in text domain. However, the use of attention-based models for video under…

cs.CV2020

Cross-modal Learning for Multi-modal Video Categorization

Palash Goyal, Saurabh Sahu, Shalini Ghosh +1

Multi-modal machine learning (ML) models can process data in multiple modalities (e.g., video, audio, text) and are useful for video content analysis in a variety of problems (e.g.…

cs.CV2020

Exploiting Temporal Coherence for Multi-modal Video Categorization

Palash Goyal, Saurabh Sahu, Shalini Ghosh +1

Multimodal ML models can process data in multiple modalities (e.g., video, images, audio, text) and are useful for video content analysis in a variety of problems (e.g., object det…

cs.LG2019

Modeling Feature Representations for Affective Speech using Generative Adversarial Networks

Saurabh Sahu, Rahul Gupta, Carol Espy-Wilson

Emotion recognition is a classic field of research with a typical setup extracting features and feeding them through a classifier for prediction. On the other hand, generative mode…