1 citations · 1 across the 3 of their papers we have counts for
9 papers
Leveraging Local Temporal Information for Multimodal Scene Classification
Saurabh Sahu, Palash Goyal
Robust video scene classification models should capture the spatial (pixel-wise) and temporal (frame-wise) characteristics of a video effectively. Transformer models with self-atte…
Can't Fool Me: Adversarially Robust Transformer for Video Understanding
Divya Choudhary, Palash Goyal, Saurabh Sahu
Deep neural networks have been shown to perform poorly on adversarial examples. To address this, several techniques have been proposed to increase robustness of a model for image c…
Enhancing Transformer for Video Understanding Using Gated Multi-Level Attention and Temporal Adversarial Training
Saurabh Sahu, Palash Goyal
The introduction of Transformer model has led to tremendous advancements in sequence modeling, especially in text domain. However, the use of attention-based models for video under…
Cross-modal Learning for Multi-modal Video Categorization
Palash Goyal, Saurabh Sahu, Shalini Ghosh +1
Multi-modal machine learning (ML) models can process data in multiple modalities (e.g., video, audio, text) and are useful for video content analysis in a variety of problems (e.g.…
Exploiting Temporal Coherence for Multi-modal Video Categorization
Palash Goyal, Saurabh Sahu, Shalini Ghosh +1
Multimodal ML models can process data in multiple modalities (e.g., video, images, audio, text) and are useful for video content analysis in a variety of problems (e.g., object det…
Modeling Feature Representations for Affective Speech using Generative Adversarial Networks
Saurabh Sahu, Rahul Gupta, Carol Espy-Wilson
Emotion recognition is a classic field of research with a typical setup extracting features and feeding them through a classifier for prediction. On the other hand, generative mode…