From the 1 of 7 linked papers with an AI index.
7 papers
CosFly-VLA: A Spatially Aware Vision-Language-Action Model for UAV Tracking
Ruilong Ren, Songsheng Cheng, Yunpeng Zhou +9
The paper introduces CosFly-VLA, a vision-language-action model for UAVs that jointly grounds target location, predicts visibility, and generates flight actions, using spatial pret…
Speculative Pipeline Decoding: Higher-Accuracy Drafting with Hidden Latency via Pipeline Parallelism
Yijiong Yu, Huazheng Wang, Shuai Yuan +2
Speculative Decoding (SD) accelerates low-concurrency LLM inference with a draft-then-verify paradigm. Mainstream methods, however, rely on multi-token prediction, which incurs com…
CosFly-Track: A Large-Scale Multi-Modal Dataset for UAV Visual Tracking via Multi-Constraint Trajectory Optimization
Xiangyue Wang, Hanxuan Chen, Songsheng Cheng +7
Recent aerial vision-language navigation (VLN) datasets have grown rapidly, but they primarily address goal-oriented navigation to static destinations, leaving UAV visual tracking…
CosFly: Plan in the Matrix, Fly in the World
Hanxuan Chen, Xiangyue Wang, Songsheng Cheng +8
We present CosFly, a box-structured planning and multimodal simulation pipeline for aerial tracking, together with CosFly-Track, a large-scale UAV dataset for dynamic target tracki…
Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap
Hanxuan Chen, Jie Zheng, Siqi Yang +9
Vision-and-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) represents a pivotal challenge in embodied artificial intelligence, focused on enabling UAVs to interpret high…
TeleEgo: Benchmarking Egocentric AI Assistants in the Wild
Jiaqi Yan, Ruilong Ren, Jingren Liu +13
Egocentric AI assistants in real-world settings must process multi-modal inputs (video, audio, text), respond in real time, and retain evolving long-term memory. However, existing…