7 papers
ProRe: A Proactive Reward System for GUI Agents via Reasoner-Actor Collaboration
Gaole Dai, Shiqi Jiang, Ting Cao +5
Reward is critical to the evaluation and training of large language models (LLMs). However, existing rule-based or model-based reward methods struggle to generalize to GUI agents,…
UniG2U-Bench: Do Unified Models Advance Multimodal Understanding?
Zimo Wen, Boxiu Li, Wanbo Zhang +11
Unified multimodal models have recently demonstrated strong generative capabilities, yet whether and when generation improves understanding remains unclear. Existing benchmarks lac…
OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation
Heyu Guo, Shanmu Wang, Ruichun Ma +5
Vision-language-action (VLA) models have shown strong generalization for robotic action prediction through large-scale vision-language pretraining. However, most existing models re…
Advancing Mobile GUI Agents: A Verifier-Driven Approach to Practical Deployment
Gaole Dai, Shiqi Jiang, Ting Cao +5
We propose V-Droid, a mobile GUI task automation agent. Unlike previous mobile agents that utilize Large Language Models (LLMs) as generators to directly generate actions at each s…
AVA: Towards Agentic Video Analytics with Vision Language Models
Yuxuan Yan, Shiqi Jiang, Ting Cao +5
AI-driven video analytics has become increasingly important across diverse domains. However, existing systems are often constrained to specific, predefined tasks, limiting their ad…
Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning
Chendong Wang, Donglin Bai, Yifan Yang +11
We present \emph{Video-in-the-Loop} (ViTL), a two-stage long-video QA framework that preserves a fixed token budget by first \emph{localizing} question-relevant interval(s) with a…