collaborators

5 papers

cs.CV2026

DynaPix: Can Vision-Language Models Identify the Exact Future?

Thong Nguyen, Vinh-Hien Do, Quynh Vo +2

Acting in a physical scene requires knowing its real later state, not a plausible one. Current evaluations often accept words or a realistic-looking image, so the predicted state i…

cs.CV2026

Predict, Then Retrieve: Cross-Instance Future-State Retrieval from Video Prefixes

Quynh Vo, Thong Nguyen, Vinh-Hien Do +2

We introduce Predictive State Retrieval (PSR), a task in which a model observes a short video prefix and a temporal question about an object's future state, then retrieves instance…

cs.CV2026

When Depth Is Better Told Than Shown: Depth-Ordinal Prompting for Vision-Language Spatial Reasoning

Quynh Vo, Phuc Dao, Cong-Duy Nguyen +1

Vision-language models (VLMs) are expected to reason about physical space -- which object is closer, what lies behind what, and how objects are arranged in 3D -- yet they still str…

cs.CL2026

TIGER: Text-Conditioned Visual Gated Routing with Acceptance Alignment for Multimodal Speculative Decoding

Quynh Vo, Cong-Duy Nguyen, Ponhvoan Srey +2

Speculative decoding accelerates autoregressive generation by letting a lightweight drafter propose multiple tokens that are verified by a larger target model. Although effective f…

cs.CL2026

ViHERMES: A Graph-Grounded Multihop Question Answering Benchmark and System for Vietnamese Healthcare Regulations

Long S. T. Nguyen, Quan M. Bui, Tin T. Ngo +3

Question Answering (QA) over regulatory documents is inherently challenging due to the need for multihop reasoning across legally interdependent texts, a requirement that is partic…