8 papers
D-VLC: Decentralized Vision-Language Collaboration for Heterogeneous Embodied Multi-Robot Systems in Unknown Environments
Yuan Zhou, Ruitong Lin, Shen Wang +6
Multi-robot systems, particularly heterogeneous robot swarms, can improve the efficiency of complex task execution through parallel collaboration and complementary capabilities. Ho…
A Kinetic Energy Perspective of Flow Matching
Ziyun Li, Huancheng Hu, Soon Hoe Lim +6
Flow-based generative models can be viewed through a physics lens: sampling transports a particle from noise to data by integrating a learned velocity field, and each sample corres…
FlyMirage: A Fully Automated Generation Pipeline for Diverse and Scalable UAV Flight Data via Generative World Model
Jinhan Li, Xijie Huang, Zhaoqi Wang +7
In the field of Vision-Language Navigation (VLN), aerial datasets remain limited in their ability to combine scale, diversity, and realism, often relying on either costly real-worl…
Vector sketch animation generation with differentiable motion trajectories
Xinding Zhu, Xinye Yang, Shuyang Zheng +4
Sketching is a direct and inexpensive means of visual expression. Though image-based sketching has been well studied, video-based sketch animation generation is still very challeng…
NavDreamer: Video Models as Zero-Shot 3D Navigators
Xijie Huang, Weiqi Gai, Tianyue Wu +5
Previous Vision-Language-Action models face critical limitations in navigation: scarce, diverse data from labor-intensive collection and static representations that fail to capture…
USS-Nav: Unified Spatio-Semantic Scene Graph for Lightweight UAV Zero-Shot Object Navigation
Weiqi Gai, Yuman Gao, Yuan Zhou +6
Zero-Shot Object Navigation in unknown environments poses significant challenges for Unmanned Aerial Vehicles (UAVs) due to the conflict between high-level semantic reasoning requi…