14 papers
DynaPix: Can Vision-Language Models Identify the Exact Future?
Thong Nguyen, Vinh-Hien Do, Quynh Vo +2
Acting in a physical scene requires knowing its real later state, not a plausible one. Current evaluations often accept words or a realistic-looking image, so the predicted state i…
When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift
Khoi Le, Tri Cao, Phong Nguyen +5
Weak-to-strong (W2S) generalization is a promising framework for scalable oversight, yet existing evaluations often test students under matched train-test distributions. Therefore,…
Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning
Lin Li, Jiawei Huang, Qihao Quan +7
In this paper, we propose the first VL gentic easoning framework for few-hot multimo…
READ: Recurrent Adapter with Partial Video-Language Alignment for Parameter-Efficient Transfer Learning in Low-Resource Video-Language Modeling
Thong Nguyen, Xiaobao Wu, Xinshuai Dong +5
Fully fine-tuning pretrained large-scale transformer models has become a popular paradigm for video-language modeling tasks, such as temporal language grounding and video-language…
DemaFormer: Damped Exponential Moving Average Transformer with Energy-Based Modeling for Temporal Language Grounding
Thong Nguyen, Xiaobao Wu, Xinshuai Dong +3
Temporal Language Grounding seeks to localize video moments that semantically correspond to a natural language query. Recent advances employ the attention mechanism to learn the re…
Perception Without Engagement: Dissecting the Causal Discovery Deficit in LMMs
Jiafeng Liang, Zhihao Zhu, Zihan Zhang +7
Although Large Multimodal Models (LMMs) have achieved strong performance on general video understanding, their susceptibility to textual prior shortcuts during causal discovery has…