latent space interventions 1model robustness 1representation analysis 1vision-language models 1visual grounding 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.CV2026
Prior Directions: Why GUI Grounding Gets Locked in the Past
Weile Gong, Zijian Lu, Mingcai Chen +3
The paper investigates how vision-language models can become locked onto outdated textual priors, causing incorrect visual grounding, and identifies recurring latent directions—cal…
cs.CV2026
Geometric Risk Control for Vision-Language Model OCR
Weile Gong, Yiping Zuo, Mingcai Chen +5
Vision-language models (VLMs) enable flexible generative optical character recognition (OCR), while their open-ended decoders can expose wrong but fluent text with weak visual supp…