From the 1 of 9 linked papers with an AI index.
9 papers
HyMobileAgent: Data-Environment Co-Scaling for Efficient GUI Agents
Hy Vision Team, Huawen Shen, Zhengyang Tang +20
The paper introduces HyMobileAgent, a vision-native mobile GUI agent that combines large multimodal models with a co-scaling framework for data and environments to enable precise p…
GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots
Sunqi Fan, Lingshan Chen, Runqi Yin +4
Data, as the fundamental substrate of modern intelligence, has greatly driven the development of current foundation models. Naturally, researchers aim to extend this paradigm to th…
Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction
Sunqi Fan, Qingle Liu, Runqi Yin +2
Video understanding is a fundamental capability for multimodal intelligence, and recent Multimodal Large Language Models (MLLMs) have achieved remarkable performance on Video Quest…
PhoneBuddy: Training Open Models for Agentic Phone Use
Zhengyang Tang, Xin Lai, Pengyuan Lyu +23
Phones are becoming an important execution surface for general-purpose agents, but training open models for reliable phone use remains difficult because the environment that matter…
PhoneHarness: Harnessing Phone-Use Agents through Mixed GUI, CLI, and Tool Actions
Chenxin Li, Zhengyao Fang, Zhengyang Tang +18
Phone agents are increasingly expected to complete real mobile workflows rather than merely predict the next screen action. However, much of the current mobile-agent literature sti…
PhoneWorld: Scaling Phone-Use Agent Environments
Zhengyang Tang, Yuxuan Liu, Xin Lai +21
A central bottleneck for phone-use agents is that controllable, reproducible environments covering real mobile behavior are hard to build at scale. Existing mobile-agent benchmarks…