collaborators

6 papers

cs.CV2026

Technical Report of RoboSpatial Challenge at CVPR 2026: Selective Reasoning Activation and Reference-Frame Disambiguation for Embodied Spatial Reasoning

Yuxiang Xie, Qi Lv, Jianming Xing +4

Vision-language models achieve strong general perception but often struggle with the spatial reasoning required for embodied tasks. We present RoboSpatialBrain, our submission to t…

cs.CV2026

Efficient Visual Pointing for Embodied AI:Agent-Driven Data Synthesis, Cross-Block Attention, and Iterative Correction

Zijian Hong, Qi Lv, Yuxiang Xie +4

Visual pointing maps a language instruction to pixel co ordinates, a core skill for embodied AI. We describe our PointArena 2026 solution, which achieves 77.2% overall accuracy and…

cs.CV2026

HATS: Hardness-Aware Trajectory Synthesis for GUI Agents

Rui Shao, Ruize Gao, Bin Xie +5

Graphical user interface (GUI) agents powered by large vision-language models (VLMs) have shown remarkable potential in automating digital tasks, highlighting the need for high-qua…

cs.RO2026

EnergyAction: Unimanual to Bimanual Composition with Energy-Based Models

Mingchen Song, Xiang Deng, Jie Wei +3

Recent advances in unimanual manipulation policies have achieved remarkable success across diverse robotic tasks through abundant training data and well-established model architect…

cs.RO2025

UAV-ON: A Benchmark for Open-World Object Goal Navigation with Aerial Agents

Jianqiang Xiao, Yuexuan Sun, Yixin Shao +6

Aerial navigation is a fundamental yet underexplored capability in embodied intelligence, enabling agents to operate in large-scale, unstructured environments where traditional nav…

cs.RO2025

Few-Shot Vision-Language Action-Incremental Policy Learning

Mingchen Song, Xiang Deng, Guoqiang Zhong +5

Recently, Transformer-based robotic manipulation methods utilize multi-view spatial representations and language instructions to learn robot motion trajectories by leveraging numer…