works on

From the 2 of 28 linked papers with an AI index.

activity
20242026
collaborators

28 papers

cs.CL2026

Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs

Vu Duc Anh, Nhat M. Hoang, Do Xuan Long +3

Achieving effective self-correction, where models verify and correct their own mistakes, remains a fundamental challenge for large language models (LLMs). In this work, we propose…

cs.CV2026

DynaPix: Can Vision-Language Models Identify the Exact Future?

Thong Nguyen, Vinh-Hien Do, Quynh Vo +2

Acting in a physical scene requires knowing its real later state, not a plausible one. Current evaluations often accept words or a realistic-looking image, so the predicted state i…

cs.CV2026

Predict, Then Retrieve: Cross-Instance Future-State Retrieval from Video Prefixes

Quynh Vo, Thong Nguyen, Vinh-Hien Do +2

We introduce Predictive State Retrieval (PSR), a task in which a model observes a short video prefix and a temporal question about an object's future state, then retrieves instance…

cs.CV2026

When Depth Is Better Told Than Shown: Depth-Ordinal Prompting for Vision-Language Spatial Reasoning

Quynh Vo, Phuc Dao, Cong-Duy Nguyen +1

The paper introduces Depth-Ordinal Prompting (DOP), a training‑free technique that converts monocular depth estimates into object‑level ordinal text cues, enabling vision‑language…

cs.CL2026

TIGER: Text-Conditioned Visual Gated Routing with Acceptance Alignment for Multimodal Speculative Decoding

Quynh Vo, Cong-Duy Nguyen, Ponhvoan Srey +2

The paper introduces TIGER, a framework that speeds up multimodal generation by dynamically selecting only the visual tokens relevant to the current textual context and training th…

cs.CV2026

Eulerian Motion Guidance: Robust Image Animation via Bidirectional Geometric Consistency

Thong Nguyen, Khoi M. Le, Cong-Duy Nguyen +3

Recent advancements in image animation have utilized diffusion models to breathe life into static images. However, existing controllable frameworks typically rely on Lagrangian mot…