7 papers
THYME: Temporal Hierarchical-Cyclic Interactivity Modeling for Video Scene Graphs in Aerial Footage
Trong-Thuan Nguyen, Pha Nguyen, Jackson Cothren +3
The rapid proliferation of video in applications such as autonomous driving, surveillance, and sports analytics necessitates robust methods for dynamic scene understanding. Despite…
HyperGLM: HyperGraph for Video Scene Graph Generation and Anticipation
Trong-Thuan Nguyen, Pha Nguyen, Jackson Cothren +2
Multimodal LLMs have advanced vision-language tasks but still struggle with understanding video scenes. To bridge this gap, Video Scene Graph Generation (VidSGG) has emerged to cap…
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models
Pha Nguyen, Sailik Sengupta, Girik Malik +2
The improved competence of generative models can help building multi-modal virtual assistants that leverage modalities beyond language. By observing humans performing multi-step ta…
SoGAR: Self-supervised Spatiotemporal Attention-based Social Group Activity Recognition
Naga VS Raviteja Chappa, Pha Nguyen, Alexander H Nelson +4
This paper introduces a novel approach to Social Group Activity Recognition (SoGAR) using Self-supervised Transformers network that can effectively utilize unlabeled video data. To…
CYCLO: Cyclic Graph Transformer Approach to Multi-Object Relationship Modeling in Aerial Videos
Trong-Thuan Nguyen, Pha Nguyen, Xin Li +3
Video scene graph generation (VidSGG) has emerged as a transformative approach to capturing and interpreting the intricate relationships among objects and their temporal dynamics i…
DINTR: Tracking via Diffusion-based Interpolation
Pha Nguyen, Ngan Le, Jackson Cothren +2
Object tracking is a fundamental task in computer vision, requiring the localization of objects of interest across video frames. Diffusion models have shown remarkable capabilities…