9 papers
AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents
Weikai Xu, Yunren Feng, Haoxiang Lei +6
Mobile GUI agents can operate apps through pixel perception and touch actions, making them a promising interface for collecting and improving long-horizon mobile interaction polici…
MiVE: Multiscale Vision-language features for reference-guided video Editing
Tong Wang, Meng Zou, Chengjing Wu +4
Reference-guided video editing takes a source video, a text instruction, and a reference image as inputs, requiring the model to faithfully apply the instructed edits while preserv…
How Mobile World Model Guides GUI Agents?
Weikai Xu, Kun Huang, Yunren Feng +10
Recent advances in vision-language models have enabled mobile GUI agents to perceive visual interfaces and execute user instructions, but reliable prediction of action consequences…
PhysBrain 1.0 Technical Report
Shijie Lian, Bin Yu, Xiaopeng Lin +10
Vision-language-action models have advanced rapidly, but robot trajectories alone provide limited coverage for learning broad physical understanding. PhysBrain 1.0 studies a comple…
ORCE: Order-Aware Alignment of Verbalized Confidence in Large Language Models
Chen Li, Xiaoling Hu, Songzhu Zheng +2
Large language models (LLMs) often produce answers with high certainty even when they are incorrect, making reliable confidence estimation essential for deployment in real-world sc…
RoboInter: A Holistic Intermediate Representation Suite Towards Robotic Manipulation
Hao Li, Ziqin Wang, Zi-han Ding +9
Advances in large vision-language models (VLMs) have stimulated growing interest in vision-language-action (VLA) systems for robot manipulation. However, existing manipulation data…