2 papers
cs.CV2026
DynaPix: Can Vision-Language Models Identify the Exact Future?
Thong Nguyen, Vinh-Hien Do, Quynh Vo +2
Acting in a physical scene requires knowing its real later state, not a plausible one. Current evaluations often accept words or a realistic-looking image, so the predicted state i…
cs.CV2026
Predict, Then Retrieve: Cross-Instance Future-State Retrieval from Video Prefixes
Quynh Vo, Thong Nguyen, Vinh-Hien Do +2
We introduce Predictive State Retrieval (PSR), a task in which a model observes a short video prefix and a temporal question about an object's future state, then retrieves instance…