6 papers
Training Skills Like Parameters via Self-Supervised Semantic Diffusion
Mo Li, Zixin Yin, Ting Cao +1
The paper introduces a self‑supervised framework that lets a language model acquire and store textual skills in an external library using diffusion‑style reconstruction loss, witho…
Sparse Compositional Flow Matching by geometric assembly from motion primitives
Yan Tang, Yuanbo Tang, Tingyu Cao +2
Embodied trajectories, such as the executable motion sequences of robotic manipulators, underwater vehicles, and mobile robots, are a fundamental output of embodied AI. Modern gene…
ProRe: A Proactive Reward System for GUI Agents via Reasoner-Actor Collaboration
Gaole Dai, Shiqi Jiang, Ting Cao +5
Reward is critical to the evaluation and training of large language models (LLMs). However, existing rule-based or model-based reward methods struggle to generalize to GUI agents,…
Advancing Mobile GUI Agents: A Verifier-Driven Approach to Practical Deployment
Gaole Dai, Shiqi Jiang, Ting Cao +5
We propose V-Droid, a mobile GUI task automation agent. Unlike previous mobile agents that utilize Large Language Models (LLMs) as generators to directly generate actions at each s…
Zoomer: Adaptive Image Focus Optimization for Black-box MLLM
Jiaxu Qian, Chendong Wang, Yifan Yang +16
Multimodal large language models (MLLMs) such as GPT-4o, Gemini Pro, and Claude 3.5 have enabled unified reasoning over text and visual inputs, yet they often hallucinate in real w…
Babel: A Scalable Pre-trained Model for Multi-Modal Sensing via Expandable Modality Alignment
Shenghong Dai, Shiqi Jiang, Yifan Yang +4
This paper presents Babel, the expandable modality alignment model, specially designed for multi-modal sensing. While there has been considerable work on multi-modality alignment,…