From the 1 of 83 linked papers with an AI index.
1 citations · 1 across the 38 of their papers we have counts for
7 papers · 1 filter
PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives
Kaixin Ding, Xi Chen, Minghong Cai +9
Video world models simulate future states conditioned on current observations and user actions. Recent systems have demonstrated impressive video consistency and action controllabi…
HyMobileAgent: Data-Environment Co-Scaling for Efficient GUI Agents
Hy Vision Team, Huawen Shen, Zhengyang Tang +20
The paper introduces HyMobileAgent, a vision-native mobile GUI agent that combines large multimodal models with a co-scaling framework for data and environments to enable precise p…
Towards Long-horizon Agentic Multimodal Search
Yifan Du, Zikang Liu, Jinbiao Peng +5
Multimodal deep search agents have shown great potential in solving complex tasks by iteratively collecting textual and visual evidence. However, managing the heterogeneous informa…
PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents
Zikang Liu, Junyi Li, Wayne Xin Zhao +3
Graphical User Interface (GUI) agents powered by Multimodal Large Language Models (MLLMs) promise human-like interaction with software applications, yet long-horizon tasks remain c…
Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search
Xin Lai, Junyi Li, Wei Li +3
Recent advances in large multimodal models have leveraged image-based tools with reinforcement learning to tackle visual problems. However, existing open-source approaches often ex…
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
Senqiao Yang, Junyi Li, Xin Lai +3
Recent advancements in vision-language models (VLMs) have improved performance by increasing the number of visual tokens, which are often significantly longer than text tokens. How…