17 papers
G-ReAct: Graph-Guided Deep Search via Structure-State Co-Evolution
Shaoxiong Yang, Mengyuan Zhang, Shaojun Lin +4
Deep search has become a fundamental capability of large language models (LLMs) for solving open-domain complex tasks. However, existing approaches typically rely on linear sequent…
SEE: Structure-aware Exploring & Exploiting for Long-horizon GUI Agent Trajectory Synthesis
Zhuohang Fan, Beichen Zhang, Yuanfa Li +4
Graphical User Interface (GUI) agents powered by vision-language models hold promise for automating real-world mobile tasks. However, progress is limited by the lack of high-covera…
ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental Synthesis
Guohong Liu, Jialei Ye, Pengzhi Gao +4
GUI agents powered by large language models are advancing rapidly, creating urgent needs for evaluation and training based on realistic environments. However, directly doing so in…
Scaling, Benchmarking, and Reasoning of Vision-Language Agents for Mobile GUI Navigation
Heng Qu, Yike Liu, Renren Jin +4
Vision-Language Models (VLMs) have shown rapid progress in mobile GUI navigation. This paper presents a systematic study of data scaling, benchmarking, and reasoning for VLM-based…
InterSketch: An Interleaved Reasoning Model with Self-correcting Visual Sketch and Stepwise Reward
Zhiwei Ning, Wenwen Tong, Xiangli Kong +12
While vision-language models (VLMs) have exhibited multi-turn visual reasoning capabilities, their reasoning trajectories remain relatively shallow and are dominated by a text-cent…
The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
MiniMax, :, Aili Chen +219
We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The…