10 papers
WorldFly: A World-Model-Based Vision-Language-Action Model for UAV Navigation
Shengtao Zheng, Kai Li, Weichen Zhang +5
End-to-end Vision-Language-Action (VLA) models have shown promise in UAV navigation. However, existing approaches typically rely on historical observations to directly predict acti…
The Point, the Vision and the Text: Does Point Cloud Boost Spatial Reasoning of Large Language Models? A Bias-Controlled Study
Weichen Zhang, Ruiying Peng, Xin Zeng +9
3D Large Language Models (LLMs) leveraging spatial information in point clouds for 3D spatial reasoning attract great attention. Despite some promising results, the advantages of p…
EventPrune: Cascaded Event-Assisted Token Pruning for Efficient First-Person Dynamic Spatial Reasoning
Pengtao Ma, Ziliang Zhou, Ciyu Ruan +7
First-person dynamic spatial reasoning requires models to track continuous motion and precise geometric structure, but the quadratic attention cost of Transformer-based Video-LLMs…
SCOPE: Skeleton Graph-Based Computation-Efficient Framework for Autonomous UAV Exploration
Kai Li, Shengtao Zheng, Linkun Xiu +4
Autonomous exploration in unknown environments is key for mobile robots, helping them perceive, map, and make decisions in complex areas. However, current methods often rely on fre…
Progressive Supernet Training for Efficient Visual Autoregressive Modeling
Xiaoyue Chen, Yuling Shi, Kaiyuan Li +5
Visual Auto-Regressive (VAR) models significantly reduce inference steps through the "next-scale" prediction paradigm. However, progressive multi-scale generation incurs substantia…
Balanced Token Pruning: Accelerating Vision Language Models Beyond Local Optimization
Kaiyuan Li, Xiaoyue Chen, Chen Gao +2
Large Vision-Language Models (LVLMs) have shown impressive performance across multi-modal tasks by encoding images into thousands of tokens. However, the large number of image toke…