7 papers
Eulerian Motion Guidance: Robust Image Animation via Bidirectional Geometric Consistency
Thong Nguyen, Khoi M. Le, Cong-Duy Nguyen +3
Recent advancements in image animation have utilized diffusion models to breathe life into static images. However, existing controllable frameworks typically rely on Lagrangian mot…
WINDQuant: Weight-Informed Neural Decision-Making for Global Mixed-Precision LLM Quantization
Phong Nam Huu Nguyen, Khoi M. Le, Cong-Duy T Nguyen +3
Quantization is an effective approach to reduce the memory footprint and inference cost of large language models (LLMs), yet maintaining performance in the ultra-low-bit regime rem…
When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift
Khoi Le, Tri Cao, Phong Nguyen +5
Weak-to-strong (W2S) generalization is a promising framework for scalable oversight, yet existing evaluations often test students under matched train-test distributions. Therefore,…
WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections
Tri Cao, Yulin Chen, Hieu Cao +8
Web agents can autonomously complete online tasks by interacting with websites, but their exposure to open web environments makes them vulnerable to prompt injection attacks embedd…
READ: Recurrent Adapter with Partial Video-Language Alignment for Parameter-Efficient Transfer Learning in Low-Resource Video-Language Modeling
Thong Nguyen, Xiaobao Wu, Xinshuai Dong +5
Fully fine-tuning pretrained large-scale transformer models has become a popular paradigm for video-language modeling tasks, such as temporal language grounding and video-language…
Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models
Tri Cao, Khoi Le, Thong Nguyen +7
While multimodal large language models (MLLMs) have advanced video understanding, they remain highly prone to hallucinations in dynamic scenes. We argue this stems from a failure i…