activity
20192025
most citedA Spatio-temporal Attention-based Model for Infant Movement Assessment from Videos

53 citations · 79 across the 12 of their papers we have counts for

collaborators
Showing cs.CVShow all

18 papers · 1 filter

cs.CV2025

Finding the Trigger: Causal Abductive Reasoning on Video Events

Thao Minh Le, Vuong Le, Kien Do +3

This paper introduces a new problem, Causal Abductive Reasoning on Video Events (CARVE), which involves identifying causal relationships between events in a video and generating hy…

cs.CV2024

SADL: An Effective In-Context Learning Method for Compositional Visual QA

Long Hoang Dang, Thao Minh Le, Vuong Le +2

Large vision-language models (LVLMs) offer a novel capability for performing in-context learning (ICL) in Visual QA. When prompted with a few demonstrations of image-question-answe…

cs.CV2023

Persistent-Transient Duality: A Multi-mechanism Approach for Modeling Human-Object Interaction

Hung Tran, Vuong Le, Svetha Venkatesh +1

Humans are highly adaptable, swiftly switching between different modes to progressively handle different tasks, situations and contexts. In Human-object interaction (HOI) activitie…

cs.CV2022

Guiding Visual Question Answering with Attention Priors

Thao Minh Le, Vuong Le, Sunil Gupta +2

The current success of modern visual reasoning systems is arguably attributed to cross-modality attention mechanisms. However, in deliberative reasoning such as in VQA, attention i…

cs.CV2022

Persistent-Transient Duality in Human Behavior Modeling

Hung Tran, Vuong Le, Svetha Venkatesh +1

We propose to model the persistent-transient duality in human behavior using a parent-child multi-channel neural network, which features a parent persistent channel that manages th…

cs.CV20216 cited

Hierarchical Object-oriented Spatio-Temporal Reasoning for Video Question Answering

Long Hoang Dang, Thao Minh Le, Vuong Le +1

Video Question Answering (Video QA) is a powerful testbed to develop new AI capabilities. This task necessitates learning to reason about objects, relations, and events across visu…