works on

From the 1 of 12 linked papers with an AI index.

collaborators

12 papers

cs.RO2026

FlowPilot: Real-Time World-Action Modeling for Agile UAV Navigation

Runqing Wang, Ding Yu, Pengyuan Min +6

We present FlowPilot, a compact world-action model for real-time onboard UAV navigation from depth. Unlike map-then-optimize pipelines that require local reconstruction or end-to-e…

cs.RO2026

AeroAct: Action-Centered World-Action Models for Language-Conditioned Quadrotor Flight

Xinhong Zhang, Qiyuan Zhu, Yubo Huang +8

The paper introduces AeroAct, a world-action model that predicts quadrotor flight actions from egocentric video, proprioceptive data, and language commands, using a video diffusion…

cs.RO2026

Asymmetric physics enables efficient learning in quadrupedal robot swarms

Yuang Zhang, Yunlong Song, Zhihao He +7

Animal collectives navigate cluttered environments through local coordination, yet robot swarms still struggle to reproduce this capability in the physical world. End-to-end learni…

cs.RO2026

Autonomous FPV Flight with Translational Optical Flow and Uncertainty Mask

Yang Deng, Yu Hu, Feng Yu +2

Autonomous FPV quadrotor flight in complex environments using a monocular RGB camera as the sole exteroceptive sensor remains a fundamental challenge. Recent research has shown tha…

cs.CV2026

PrecisionCUA: Iterative Visual Refinement for Pixel-Precise Cursor Grounding in Code Editors

Himangi Mittal, Gaurav Mittal, Nelson Daniel Troncoso +1

Computer Use Agents (CUAs) fundamentally rely on graphical user interface (GUI) grounding to translate language instructions into executable screen actions, but editing-level groun…

cs.RO2026

E2E-Fly: An Integrated Training-to-Deployment System for End-to-End Quadrotor Autonomy

Fangyu Sun, Fanxing Li, Linzuo Zhang +5

Training and transferring learning-based policies for quadrotors from simulation to reality remains challenging due to inefficient visual rendering, physical modeling inaccuracies,…