5 papers
Finding the Trigger: Causal Abductive Reasoning on Video Events
Thao Minh Le, Vuong Le, Kien Do +3
This paper introduces a new problem, Causal Abductive Reasoning on Video Events (CARVE), which involves identifying causal relationships between events in a video and generating hy…
Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models
Quang-Hung Le, Long Hoang Dang, Ngan Le +2
Existing Large Vision-Language Models (LVLMs) excel at matching concepts across multi-modal inputs but struggle with compositional concepts and high-level relationships between ent…
SADL: An Effective In-Context Learning Method for Compositional Visual QA
Long Hoang Dang, Thao Minh Le, Vuong Le +2
Large vision-language models (LVLMs) offer a novel capability for performing in-context learning (ICL) in Visual QA. When prompted with a few demonstrations of image-question-answe…
Revisiting the Dataset Bias Problem from a Statistical Perspective
Kien Do, Dung Nguyen, Hung Le +6
In this paper, we study the "dataset bias" problem from a statistical standpoint, and identify the main cause of the problem as the strong correlation between a class attribute u a…
Video Dialog as Conversation about Objects Living in Space-Time
Hoang-Anh Pham, Thao Minh Le, Vuong Le +2
It would be a technological feat to be able to create a system that can hold a meaningful conversation with humans about what they watch. A setup toward that goal is presented as a…