3 papers
cs.CV2026
How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
Weiming Li, Yan Shao, Jing Yang +6
Graphical user interface (GUI) grounding is a fundamental task for building GUI agents. However, general vision-language models (VLMs) struggle with this task due to a lack of spec…
cs.CL2026
TRE: Encouraging Exploration in the Trust Region
Chao Huang, Yujing Lu, Quangang Li +8
Entropy regularization is a standard technique in reinforcement learning (RL) to enhance exploration, yet it yields negligible effects or even degrades performance in Large Languag…
cs.CL2026
DomainCQA: Crafting Knowledge-Intensive QA from Domain-Specific Charts
Yujing Lu, Ling Zhong, Jing Yang +5
Chart Question Answering (CQA) evaluates Multimodal Large Language Models (MLLMs) on visual understanding and reasoning over chart data. However, existing benchmarks mostly test su…