36 citations · 71 across the 15 of their papers we have counts for
20 papers
Mamba Fusion: Learning Actions Through Questioning
Zhikang Dong, Apoorva Beedu, Jason Sheinkopf +1
Video Language Models (VLMs) are crucial for generalizing across diverse tasks and using language cues to enhance learning. While transformer-based architectures have been the de f…
Limitations in Employing Natural Language Supervision for Sensor-Based Human Activity Recognition -- And Ways to Overcome Them
Harish Haresamudram, Apoorva Beedu, Mashfiqui Rabbi +3
Cross-modal contrastive pre-training between natural language and other modalities, e.g., vision and audio, has demonstrated astonishing performance and effectiveness across a dive…
Investigating Enhancements to Contrastive Predictive Coding for Human Activity Recognition
Harish Haresamudram, Irfan Essa, Thomas Ploetz
The dichotomy between the challenging nature of obtaining annotations for activities, and the more straightforward nature of data collection from wearables, has resulted in signifi…
Multi-Stage Based Feature Fusion of Multi-Modal Data for Human Activity Recognition
Hyeongju Choi, Apoorva Beedu, Harish Haresamudram +1
To properly assist humans in their needs, human activity recognition (HAR) systems need the ability to fuse information from multiple modalities. Our hypothesis is that multimodal…
End-to-End Multimodal Representation Learning for Video Dialog
Huda Alamri, Anthony Bilic, Michael Hu +2
Video-based dialog task is a challenging multimodal learning task that has received increasing attention over the past few years with state-of-the-art obtaining new performance rec…
Finding Islands of Predictability in Action Forecasting
Daniel Scarafoni, Irfan Essa, Thomas Ploetz
We address dense action forecasting: the problem of predicting future action sequence over long durations based on partial observation. Our key insight is that future action sequen…