4 papers
ProRe: A Proactive Reward System for GUI Agents via Reasoner-Actor Collaboration
Gaole Dai, Shiqi Jiang, Ting Cao +5
Reward is critical to the evaluation and training of large language models (LLMs). However, existing rule-based or model-based reward methods struggle to generalize to GUI agents,…
Advancing Mobile GUI Agents: A Verifier-Driven Approach to Practical Deployment
Gaole Dai, Shiqi Jiang, Ting Cao +5
We propose V-Droid, a mobile GUI task automation agent. Unlike previous mobile agents that utilize Large Language Models (LLMs) as generators to directly generate actions at each s…
Diffusion-Aided Bandwidth-Efficient Semantic Communication with Adaptive Requests
Xuesong Wang, Xinyan Xie, Mo Li +1
Semantic communication focuses on conveying the task-relevant meaning rather than exact bitwise recovery. For image transmission with a generative receiver, relying only on text de…
Babel: A Scalable Pre-trained Model for Multi-Modal Sensing via Expandable Modality Alignment
Shenghong Dai, Shiqi Jiang, Yifan Yang +4
This paper presents Babel, the expandable modality alignment model, specially designed for multi-modal sensing. While there has been considerable work on multi-modality alignment,…