17 papers
MentalThink: Shaping Thoughts in Mental SVG World
Kangheng Lin, Jisheng Yin, Dingming Li +11
We introduce MentalThink, a visual-symbolic reasoning paradigm that equips Multimodal LLMs (MLLMs) with an executable mechanism for "mental" visualization. The core of MentalThink…
PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception
Yana Wei, Hongbo Peng, Yanlin Lai +14
We introduce PerceptionRubrics, a rubric-based evaluation framework that addresses the gap between saturated benchmark scores and real-world brittleness. Shifting evaluation from h…
WebVR: Benchmarking Multimodal LLMs for WebPage Recreation from Videos via Human-Aligned Visual Rubrics
Yuhong Dai, Yanlin Lai, Mitt Huang +9
Existing web-generation benchmarks rely on text prompts or static screenshots as input. However, videos naturally convey richer signals such as interaction flow, transition timing,…
Step-GUI Technical Report
Haolong Yan, Jia Wang, Xin Huang +95
Recent advances in multimodal large language models unlock unprecedented opportunities for GUI automation. However, a fundamental challenge remains: how to efficiently acquire high…
GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning
Haolong Yan, Yeqing Shen, Xin Huang +9
With the rapid development of Large Vision Language Models, the focus of Graphical User Interface (GUI) agent tasks shifts from single-screen tasks to complex screen navigation cha…
Thinking by Doing: Building Efficient World Model Reasoning in LLMs via Multi-turn Interaction
Bao Shu, Yan Cai, Jianjian Sun +11
Developing robust world model reasoning is crucial for large language model (LLM) agents to plan and interact in complex environments. While multi-turn interaction offers a superio…