2 papers
cs.CV2025
More Than the Final Answer: Improving Visual Extraction and Logical Consistency in Vision-Language Models
Hoang Anh Just, Yifei Fan, Handong Zhao +6
Reinforcement learning from verifiable rewards (RLVR) has recently been extended from text-only LLMs to vision-language models (VLMs) to elicit long-chain multimodal reasoning. How…
cs.CL2025
GUI-Bee: Align GUI Action Grounding to Novel Environments via Autonomous Exploration
Yue Fan, Handong Zhao, Ruiyi Zhang +3
Graphical User Interface (GUI) action grounding is a critical step in GUI automation that maps language instructions to actionable elements on GUI screens. Most recent works of GUI…