From the 1 of 10 linked papers with an AI index.
10 papers
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
Hanzhang Zhou, Panrong Tong, Xu Zhang +13
The paper introduces Qwen-UI-Agent, a foundation model for GUI agents that can operate across mobile, desktop, web, and search environments, combining GUI actions with CLI commands…
One Forward Beats Two: InnerZoom for Accurate and Efficient GUI Grounding
Chen Liu, Ling Chen, Hanzhang Zhou +5
MLLM-based GUI grounding methods commonly formulate target localization as autoregressive coordinate generation, enabling models to leverage the strong instruction-following and se…
ViSurf: Visual Supervised-and-Reinforcement Fine-Tuning for Large Vision-and-Language Models
Yuqi Liu, Liangyu Chen, Jiazhen Liu +4
Post-training Large Vision-and-Language Models (LVLMs) typically involves Supervised Fine-Tuning (SFT) for knowledge injection or Reinforcement Learning with Verifiable Rewards (RL…
MCSC-Bench: Multimodal Context-to-Script Creation for Realistic Video Production
Huanran Hu, Zihui Ren, Dingyi Yang +4
Real-world video creation often involves a complex reasoning workflow of selecting relevant shots from noisy materials, planning missing shots for narrative completeness, and organ…
POLYCHARTQA: Benchmarking Large Vision-Language Models with Multilingual Chart Question Answering
Yichen Xu, Liangyu Chen, Liang Zhang +3
Charts are a universally adopted medium for data communication, yet existing chart understanding benchmarks are overwhelmingly English-centric, limiting their accessibility and rel…
MobileWorld: Benchmarking Autonomous Mobile Agents in Agent-User Interactive and MCP-Augmented Environments
Quyu Kong, Xu Zhang, Zhenyu Yang +10
Among existing online mobile-use benchmarks, AndroidWorld has emerged as the dominant benchmark due to its reproducible environment and deterministic evaluation; however, recent ag…