6 papers
Adaptive Residual-Update Steering for Low-Overhead Hallucination Mitigation in Large Vision Language Models
Zhengtao Zou, Ya Gao, Jiarui Guan +2
Large Vision-Language Models (LVLMs) typically process visual inputs as a prefix to the language decoder. As the model autoregressively generates text, this initial visual informat…
DTP: A Simple yet Effective Distracting Token Pruning Framework for Vision-Language Action Models
Chenyang Li, Jieyuan Liu, Bin Li +5
Vision-Language Action (VLA) models have shown remarkable progress in robotic manipulation by leveraging the powerful perception abilities of Vision-Language Models (VLMs) to under…
See the Forest and the Trees: A Synergistic Reasoning Framework for Knowledge-Based Visual Question Answering
Junjie Wang, Yunhan Tang, Yijie Wang +4
Multimodal Large Language Models (MLLMs) have pushed the frontiers of Knowledge-Based Visual Question Answering (KBVQA), yet their reasoning is fundamentally bottlenecked by a reli…
Low-Cost Test-Time Adaptation for Robust Video Editing
Jianhui Wang, Yinda Chen, Yangfan He +6
Video editing is a critical component of content creation that transforms raw footage into coherent works aligned with specific visual and narrative objectives. Existing approaches…
XeMap: Contextual Referring in Large-Scale Remote Sensing Environments
Yuxi Li, Lu Si, Yujie Hou +4
Advancements in remote sensing (RS) imagery have provided high-resolution detail and vast coverage, yet existing methods, such as image-level captioning/retrieval and object-level…
EffOWT: Transfer Visual Language Models to Open-World Tracking Efficiently and Effectively
Bingyang Wang, Kaer Huang, Bin Li +4
Open-World Tracking (OWT) aims to track every object of any category, which requires the model to have strong generalization capabilities. Trackers can improve their generalization…