36 citations · 188 across the 35 of their papers we have counts for
11 papers · 1 filter
Cascaded Compositional Residual Learning for Complex Interactive Behaviors
K. Niranjan Kumar, Irfan Essa, Sehoon Ha
Real-world autonomous missions often require rich interaction with nearby objects, such as doors or switches, along with effective navigation. However, such complex behaviors are d…
MAGVIT: Masked Generative Video Transformer
Lijun Yu, Yong Cheng, Kihyuk Sohn +8
We introduce the MAsked Generative VIdeo Transformer, MAGVIT, to tackle various video synthesis tasks with a single model. We introduce a 3D tokenizer to quantize a video into spat…
Investigating Enhancements to Contrastive Predictive Coding for Human Activity Recognition
Harish Haresamudram, Irfan Essa, Thomas Ploetz
The dichotomy between the challenging nature of obtaining annotations for activities, and the more straightforward nature of data collection from wearables, has resulted in signifi…
Multi-Stage Based Feature Fusion of Multi-Modal Data for Human Activity Recognition
Hyeongju Choi, Apoorva Beedu, Harish Haresamudram +1
To properly assist humans in their needs, human activity recognition (HAR) systems need the ability to fuse information from multiple modalities. Our hypothesis is that multimodal…
Video based Object 6D Pose Estimation using Transformers
Apoorva Beedu, Huda Alamri, Irfan Essa
We introduce a Transformer based 6D Object Pose Estimation framework VideoPose, comprising an end-to-end attention based modelling architecture, that attends to previous frames in…
End-to-End Multimodal Representation Learning for Video Dialog
Huda Alamri, Anthony Bilic, Michael Hu +2
Video-based dialog task is a challenging multimodal learning task that has received increasing attention over the past few years with state-of-the-art obtaining new performance rec…