4 papers
PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud
Chenghua Wang, Daliang Xu, Dongqi Cai +24
Physical AI policies require inference throughout their lifecycle, including model evaluation, cloud reinforcement learning rollout, edge GPU serving, and onboard deployment. Altho…
MobileViews: A Million-scale and Diverse Mobile GUI Dataset
Longxi Gao, Li Zhang, Shihe Wang +6
Visual language models (VLMs) empower mobile GUI agents to interpret complex mobile screens and respond to user requests. Training such capable agents requires large-scale, high-qu…
GUI-Shift: Enhancing VLM-Based GUI Agents through Self-supervised Reinforcement Learning
Longxi Gao, Li Zhang, Pengzhi Gao +3
Training effective Vision-Language Models (VLMs) for GUI agents typically depends on large-scale annotated datasets, whose collection is both labor-intensive and error-prone. We in…
Does Chain-of-Thought Reasoning Help Mobile GUI Agent? An Empirical Study
Li Zhang, Longxi Gao, Mengwei Xu
Reasoning capabilities have significantly improved the performance of vision-language models (VLMs) in domains such as mathematical problem-solving, coding, and visual question-ans…