4 citations · 8 across the 7 of their papers we have counts for
7 papers · 1 filter
HierSum: A Global and Local Attention Mechanism for Video Summarization
Apoorva Beedu, Irfan Essa
Video summarization creates an abridged version (i.e., a summary) that provides a quick overview of the video while retaining pertinent information. In this work, we focus on summa…
Exploring Efficient Foundational Multi-modal Models for Video Summarization
Karan Samel, Apoorva Beedu, Nitish Sontakke +1
Foundational models are able to generate text outputs given prompt instructions and text, audio, or image inputs. Recently these models have been combined to perform tasks on video…
Mamba Fusion: Learning Actions Through Questioning
Zhikang Dong, Apoorva Beedu, Jason Sheinkopf +1
Video Language Models (VLMs) are crucial for generalizing across diverse tasks and using language cues to enhance learning. While transformer-based architectures have been the de f…
On the Efficacy of Text-Based Input Modalities for Action Anticipation
Apoorva Beedu, Harish Haresamudram, Karan Samel +1
Anticipating future actions is a highly challenging task due to the diversity and scale of potential future actions; yet, information from different modalities help narrow down pla…
Multimodal Contrastive Learning with Hard Negative Sampling for Human Activity Recognition
Hyeongju Choi, Apoorva Beedu, Irfan Essa
Human Activity Recognition (HAR) systems have been extensively studied by the vision and ubiquitous computing communities due to their practical applications in daily life, such as…
Multi-Stage Based Feature Fusion of Multi-Modal Data for Human Activity Recognition
Hyeongju Choi, Apoorva Beedu, Harish Haresamudram +1
To properly assist humans in their needs, human activity recognition (HAR) systems need the ability to fuse information from multiple modalities. Our hypothesis is that multimodal…