7 papers
D-VLC: Decentralized Vision-Language Collaboration for Heterogeneous Embodied Multi-Robot Systems in Unknown Environments
Yuan Zhou, Ruitong Lin, Shen Wang +6
Multi-robot systems, particularly heterogeneous robot swarms, can improve the efficiency of complex task execution through parallel collaboration and complementary capabilities. Ho…
NavDreamer: Video Models as Zero-Shot 3D Navigators
Xijie Huang, Weiqi Gai, Tianyue Wu +5
Previous Vision-Language-Action models face critical limitations in navigation: scarce, diverse data from labor-intensive collection and static representations that fail to capture…
Interpretable Logical Anomaly Classification via Constraint Decomposition and Instruction Fine-Tuning
Xufei Zhang, Xinjiao Zhou, Ziling Deng +2
Logical anomalies are violations of predefined constraints on object quantity, spatial layout, and compositional relationships in industrial images. While prior work largely treats…
USS-Nav: Unified Spatio-Semantic Scene Graph for Lightweight UAV Zero-Shot Object Navigation
Weiqi Gai, Yuman Gao, Yuan Zhou +6
Zero-Shot Object Navigation in unknown environments poses significant challenges for Unmanned Aerial Vehicles (UAVs) due to the conflict between high-level semantic reasoning requi…
Flying in Clutter on Monocular RGB by Learning in 3D Radiance Fields with Domain Adaptation
Xijie Huang, Jinhan Li, Tianyue Wu +3
Modern autonomous navigation systems predominantly rely on lidar and depth cameras. However, a fundamental question remains: Can flying robots navigate in clutter using solely mono…
VLA-AN: An Efficient and Onboard Vision-Language-Action Framework for Aerial Navigation in Complex Environments
Yuze Wu, Mo Zhu, Xingxing Li +6
This paper proposes VLA-AN, an efficient and onboard Vision-Language-Action (VLA) framework dedicated to autonomous drone navigation in complex environments. VLA-AN addresses four…