activity
20192026
most citedHierarchical Object-oriented Spatio-Temporal Reasoning for Video Question Answering

6 citations · 6 across the 11 of their papers we have counts for

collaborators
Showing cs.CVShow all

13 papers · 1 filter

cs.CV2025

Planner-Refiner: Dynamic Space-Time Refinement for Vision-Language Alignment in Videos

Tuyen Tran, Thao Minh Le, Quang-Hung Le +1

Vision-language alignment in video must address the complexity of language, evolving interacting entities, their action chains, and semantic gaps between language and vision. This…

cs.CV2025

Towards Agentic AI for Multimodal-Guided Video Object Segmentation

Tuyen Tran, Thao Minh Le, Truyen Tran

Referring-based Video Object Segmentation is a multimodal problem that requires producing fine-grained segmentation results guided by external cues. Traditional approaches to this…

cs.CV2025

Finding the Trigger: Causal Abductive Reasoning on Video Events

Thao Minh Le, Vuong Le, Kien Do +3

This paper introduces a new problem, Causal Abductive Reasoning on Video Events (CARVE), which involves identifying causal relationships between events in a video and generating hy…

cs.CV2024

Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models

Quang-Hung Le, Long Hoang Dang, Ngan Le +2

Existing Large Vision-Language Models (LVLMs) excel at matching concepts across multi-modal inputs but struggle with compositional concepts and high-level relationships between ent…

cs.CV2024

Unified Framework with Consistency across Modalities for Human Activity Recognition

Tuyen Tran, Thao Minh Le, Hung Tran +1

Recognizing human activities in videos is challenging due to the spatio-temporal complexity and context-dependence of human interactions. Prior studies often rely on single input m…

cs.CV2022

Deep Neural Networks for Visual Reasoning

Thao Minh Le

Visual perception and language understanding are - fundamental components of human intelligence, enabling them to understand and reason about objects and their interactions. It is…