From the 1 of 9 linked papers with an AI index.
9 papers
CosFly-VLA: A Spatially Aware Vision-Language-Action Model for UAV Tracking
Ruilong Ren, Songsheng Cheng, Yunpeng Zhou +9
The paper introduces CosFly-VLA, a vision-language-action model for UAVs that jointly grounds target location, predicts visibility, and generates flight actions, using spatial pret…
Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms
Muyang He, Hanzhong Guo, Junxiong Lin +1
The rapid evolution of video generation has enabled models to simulate complex physical dynamics and long-horizon causalities, positioning them as potential world simulators. Howev…
Decomposing Subject-Driven Image Generation via Intermediate Structural Prediction
Hanzhong Guo, Yizhou Yu
Subject-driven text-to-image generation still struggles to preserve high-frequency identity details such as logos, patterns, and text. Existing methods typically operate directly i…
CosFly-Track: A Large-Scale Multi-Modal Dataset for UAV Visual Tracking via Multi-Constraint Trajectory Optimization
Xiangyue Wang, Hanxuan Chen, Songsheng Cheng +7
Recent aerial vision-language navigation (VLN) datasets have grown rapidly, but they primarily address goal-oriented navigation to static destinations, leaving UAV visual tracking…
Leveraging Verifier-Based Reinforcement Learning in Image Editing
Hanzhong Guo, Jie Wu, Jie Liu +6
While Reinforcement Learning from Human Feedback (RLHF) has become a pivotal paradigm for text-to-image generation, its application to image editing remains largely unexplored. A k…
CosFly: Plan in the Matrix, Fly in the World
Hanxuan Chen, Xiangyue Wang, Songsheng Cheng +8
We present CosFly, a box-structured planning and multimodal simulation pipeline for aerial tracking, together with CosFly-Track, a large-scale UAV dataset for dynamic target tracki…