activity
20212025
most citedMulti-Camera Multi-Object Tracking on the Move via Single-Stage Global Association Approach

4 citations · 7 across the 4 of their papers we have counts for

collaborators
Showing cs.CVShow all

12 papers · 1 filter

cs.CV2025

THYME: Temporal Hierarchical-Cyclic Interactivity Modeling for Video Scene Graphs in Aerial Footage

Trong-Thuan Nguyen, Pha Nguyen, Jackson Cothren +3

The rapid proliferation of video in applications such as autonomous driving, surveillance, and sports analytics necessitates robust methods for dynamic scene understanding. Despite…

cs.CV20251 cited

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models

Pha Nguyen, Sailik Sengupta, Girik Malik +2

The improved competence of generative models can help building multi-modal virtual assistants that leverage modalities beyond language. By observing humans performing multi-step ta…

cs.CV2024

HyperGLM: HyperGraph for Video Scene Graph Generation and Anticipation

Trong-Thuan Nguyen, Pha Nguyen, Jackson Cothren +2

Multimodal LLMs have advanced vision-language tasks but still struggle with understanding video scenes. To bridge this gap, Video Scene Graph Generation (VidSGG) has emerged to cap…

cs.CV2024

DINTR: Tracking via Diffusion-based Interpolation

Pha Nguyen, Ngan Le, Jackson Cothren +2

Object tracking is a fundamental task in computer vision, requiring the localization of objects of interest across video frames. Diffusion models have shown remarkable capabilities…

cs.CV2024

CYCLO: Cyclic Graph Transformer Approach to Multi-Object Relationship Modeling in Aerial Videos

Trong-Thuan Nguyen, Pha Nguyen, Xin Li +3

Video scene graph generation (VidSGG) has emerged as a transformative approach to capturing and interpreting the intricate relationships among objects and their temporal dynamics i…

cs.CV2023

HIG: Hierarchical Interlacement Graph Approach to Scene Graph Generation in Video Understanding

Trong-Thuan Nguyen, Pha Nguyen, Khoa Luu

Visual interactivity understanding within visual scenes presents a significant challenge in computer vision. Existing methods focus on complex interactivities while leveraging a si…