From the 1 of 5 linked papers with an AI index.
5 papers
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention
Yan Zhang, Yinan Wu, Haoran Duan +1
Vision-Language-Action (VLA) models have driven significant progress in robotic manipulation, yet they fundamentally struggle with the vision-override phenomenon. Driven by the sev…
Calibrate Before Reason: Robust Visual Token Reduction against Semantic Drift in VLMs
Jiasheng Li, Zhong Ji, Yan Zhang +1
The paper proposes CaRe, a training‑free method that calibrates compact visual representations before reasoning to keep semantic consistency when reducing visual tokens in large vi…
TAP-RAG: Task-Aware Policy Control for Long-Document Multimodal Question Answering
Zhong Ji, Keqi Jin, Yan Zhang +1
Long-document multimodal question answering requires more than retrieving relevant chunks from a large document. Different queries require different evidence behavior. Existing mul…
iEBAKER: Improved Remote Sensing Image-Text Retrieval Framework via Eliminate Before Align and Keyword Explicit Reasoning
Yan Zhang, Zhong Ji, Changxu Meng +2
Recent studies focus on the Remote Sensing Image-Text Retrieval (RSITR), which aims at searching for the corresponding targets based on the given query. Among these efforts, the ap…
Underlying Semantic Diffusion for Effective and Efficient In-Context Learning
Zhong Ji, Weilong Cao, Yan Zhang +3
Diffusion models has emerged as a powerful framework for tasks like image controllable generation and dense prediction. However, existing models often struggle to capture underlyin…