From the 1 of 62 linked papers with an AI index.
6 citations · 6 across the 21 of their papers we have counts for
7 papers · 1 filter
Predict, Then Retrieve: Cross-Instance Future-State Retrieval from Video Prefixes
Quynh Vo, Thong Nguyen, Vinh-Hien Do +2
We introduce Predictive State Retrieval (PSR), a task in which a model observes a short video prefix and a temporal question about an object's future state, then retrieves instance…
Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models
Tri Cao, Khoi Le, Thong Nguyen +7
While multimodal large language models (MLLMs) have advanced video understanding, they remain highly prone to hallucinations in dynamic scenes. We argue this stems from a failure i…
Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation
Thong Thanh Nguyen, Xiaobao Wu, Yi Bin +3
To equip artificial intelligence with a comprehensive understanding towards a temporal world, video and 4D panoptic scene graph generation abstracts visual data into nodes to repre…
Multi-Scale Contrastive Learning for Video Temporal Grounding
Thong Thanh Nguyen, Yi Bin, Xiaobao Wu +4
Temporal grounding, which localizes video moments related to a natural language query, is a core problem of vision-language learning and video understanding. To encode video moment…
CutPaste&Find: Efficient Multimodal Hallucination Detector with Visual-aid Knowledge Base
Cong-Duy Nguyen, Xiaobao Wu, Duc Anh Vu +3
Large Vision-Language Models (LVLMs) have demonstrated impressive multimodal reasoning capabilities, but they remain susceptible to hallucination, particularly object hallucination…
Enhancing Multimodal Entity Linking with Jaccard Distance-based Conditional Contrastive Learning and Contextual Visual Augmentation
Cong-Duy Nguyen, Xiaobao Wu, Thong Nguyen +5
Previous research on multimodal entity linking (MEL) has primarily employed contrastive learning as the primary objective. However, using the rest of the batch as negative samples…