activity
20222026
most citedMulti-Stage Based Feature Fusion of Multi-Modal Data for Human Activity Recognition

4 citations · 8 across the 7 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2025

HierSum: A Global and Local Attention Mechanism for Video Summarization

Apoorva Beedu, Irfan Essa

Video summarization creates an abridged version (i.e., a summary) that provides a quick overview of the video while retaining pertinent information. In this work, we focus on summa…

cs.CV20241 cited

Exploring Efficient Foundational Multi-modal Models for Video Summarization

Karan Samel, Apoorva Beedu, Nitish Sontakke +1

Foundational models are able to generate text outputs given prompt instructions and text, audio, or image inputs. Recently these models have been combined to perform tasks on video…

cs.CV2024

Mamba Fusion: Learning Actions Through Questioning

Zhikang Dong, Apoorva Beedu, Jason Sheinkopf +1

Video Language Models (VLMs) are crucial for generalizing across diverse tasks and using language cues to enhance learning. While transformer-based architectures have been the de f…

cs.CV2024

On the Efficacy of Text-Based Input Modalities for Action Anticipation

Apoorva Beedu, Harish Haresamudram, Karan Samel +1

Anticipating future actions is a highly challenging task due to the diversity and scale of potential future actions; yet, information from different modalities help narrow down pla…

cs.CV20232 cited

Multimodal Contrastive Learning with Hard Negative Sampling for Human Activity Recognition

Hyeongju Choi, Apoorva Beedu, Irfan Essa

Human Activity Recognition (HAR) systems have been extensively studied by the vision and ubiquitous computing communities due to their practical applications in daily life, such as…

cs.CV20224 cited

Multi-Stage Based Feature Fusion of Multi-Modal Data for Human Activity Recognition

Hyeongju Choi, Apoorva Beedu, Harish Haresamudram +1

To properly assist humans in their needs, human activity recognition (HAR) systems need the ability to fuse information from multiple modalities. Our hypothesis is that multimodal…