From the 2 of 6 linked papers with an AI index.
6 papers
DynaPix: Can Vision-Language Models Identify the Exact Future?
Thong Nguyen, Vinh-Hien Do, Quynh Vo +2
Acting in a physical scene requires knowing its real later state, not a plausible one. Current evaluations often accept words or a realistic-looking image, so the predicted state i…
Predict, Then Retrieve: Cross-Instance Future-State Retrieval from Video Prefixes
Quynh Vo, Thong Nguyen, Vinh-Hien Do +2
We introduce Predictive State Retrieval (PSR), a task in which a model observes a short video prefix and a temporal question about an object's future state, then retrieves instance…
When Depth Is Better Told Than Shown: Depth-Ordinal Prompting for Vision-Language Spatial Reasoning
Quynh Vo, Phuc Dao, Cong-Duy Nguyen +1
The paper introduces Depth-Ordinal Prompting (DOP), a training‑free technique that converts monocular depth estimates into object‑level ordinal text cues, enabling vision‑language…
TIGER: Text-Conditioned Visual Gated Routing with Acceptance Alignment for Multimodal Speculative Decoding
Quynh Vo, Cong-Duy Nguyen, Ponhvoan Srey +2
The paper introduces TIGER, a framework that speeds up multimodal generation by dynamically selecting only the visual tokens relevant to the current textual context and training th…
Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models
Tri Cao, Khoi Le, Thong Nguyen +7
While multimodal large language models (MLLMs) have advanced video understanding, they remain highly prone to hallucinations in dynamic scenes. We argue this stems from a failure i…
ViHERMES: A Graph-Grounded Multihop Question Answering Benchmark and System for Vietnamese Healthcare Regulations
Long S. T. Nguyen, Quan M. Bui, Tin T. Ngo +3
Question Answering (QA) over regulatory documents is inherently challenging due to the need for multihop reasoning across legally interdependent texts, a requirement that is partic…