3 papers
cs.CV2026
EgoSIS: From Factorized Visual Ego-Transitions to Motion-Canonical Spatial Evidence for UAV Reasoning
Jingpu Yang, Fengxian Ji, Mingxuan Cui +4
UAV video question answering requires separating camera motion from changes in the scene, but RGB-only multimodal models receive no explicit, stable reference for that separation.…
cs.CV2026
GAAT: Geometry-Aware Alignment Transformer for Multimodal UAV Perception
Jingpu Yang, Debin Tang, Yilin Sun +4
Unmanned aerial vehicle (UAV) multimodal perception integrates visible (RGB), infrared (IR), synthetic aperture radar (SAR), and depth sensors for scene understanding under diverse…
cs.CV2026
Zero-Parameter Geometric Gating for Temporally Stable Low-Altitude UAV Video Semantic Segmentation
Jingpu Yang, Fengxian Ji, Zhengzhao Lai +3
Video semantic segmentation for low-altitude UAVs requires temporal consistency, yet dense optical flow introduces spatially structured noise in the planar regions that dominate ae…