Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning
Shaojie Zhang, Pei Fu, Ruoceng Zhang +8
Autonomous graphical user interface (GUI) agents rely on accurate GUI grounding, which maps language instructions to on-screen coordinates, to execute user commands. However, curre…
cs.CV2026
Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models
Longwei Xu, Feng Feng, Shaojie Zhang +7
Optical Character Recognition (OCR) is increasingly regarded as a foundational capability for modern vision-language models (VLMs), enabling them not only to read text in images bu…