3 papers
cs.CV2026
AutoFocus: Uncertainty-Aware Active Visual Search for GUI Grounding
Ruilin Yao, Shegnwu Xiong, Tianyu Zou +2
Vision-Language Models (VLMs) have enabled autonomous GUI agents that translate natural language instructions into executable screen coordinates. However, grounding performance deg…
cs.CL2025
Aligning MLLM Benchmark With Human Preferences via Structural Equation Modeling
Shengwu. Xiong, Tianyu. Zou, Cong. Wang +1
Evaluating multimodal large language models (MLLMs) is fundamentally challenged by the absence of structured, interpretable, and theoretically grounded benchmarks; current heuristi…
cs.CV2025
MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
Peng Xu, Shengwu Xiong, Jiajun Zhang +125
This paper reviews the MARS2 2025 Challenge on Multimodal Reasoning. We aim to bring together different approaches in multimodal machine learning and LLMs via a large benchmark. We…