16 papers
Robustifying Vision-Language Models via Test-Time Prompt Adaptation
Xingyu Zhu, Huanshen Wu, Shuo Wang +4
Pre-trained Vision-Language Models (VLMs) such as CLIP achieve strong zero-shot generalization, but their performance degrades sharply under adversarial perturbations. Existing tes…
Temporal Evidence Routing with Structured Visual Evidence for TimeLogicQA
Yuyang Sun, Yongliang Wu, Xingyu Zhu +8
TimeLogicQA evaluates whether video question answering systems can reason over temporal relations such as event existence, ordering, persistence, boundary conditions, and overlap.…
Adaptive Dense Evidence Refinement for Video Relational Reasoning for VRR-QA Challenge
Yuyang Sun, Yongliang Wu, Xingyu Zhu +8
VRR-QA evaluates whether video-language systems can infer spatial, temporal, viewpoint, depth, and visibility relations that are not always resolved by a single frame. We present a…
Dual-Route Top-K Retrieval with 1v1 VLM Reranking for the CoVR-R
Yuyang Sun, Yongliang Wu, Xingyu Zhu +8
We describe \emph{Dual-Route Top-K Retrieval with 1v1 VLM Reranking} for the CoVR-R challenge. The method treats composed video retrieval as two coupled problems: finding a suffici…
Mitigating Hallucinations in Large Vision-Language Models without Performance Degradation
Xingyu Zhu, Junfeng Fang, Shuo Wang +4
Large Vision-Language Models (LVLMs) exhibit powerful generative capabilities but frequently produce hallucinations that compromise output reliability. Fine-tuning on annotated dat…
Principled Steering via Null-space Projection for Jailbreak Defense in Vision-Language Models
Xingyu Zhu, Beier Zhu, Shuo Wang +4
As vision-language models (VLMs) are increasingly deployed in open-world scenarios, they can be easily induced by visual jailbreak attacks to generate harmful content, posing serio…