2 papers
cs.CV2026
EgoSIS: From Factorized Visual Ego-Transitions to Motion-Canonical Spatial Evidence for UAV Reasoning
Jingpu Yang, Fengxian Ji, Mingxuan Cui +4
UAV video question answering requires separating camera motion from changes in the scene, but RGB-only multimodal models receive no explicit, stable reference for that separation.…
cs.CV2026
GAAT: Geometry-Aware Alignment Transformer for Multimodal UAV Perception
Jingpu Yang, Debin Tang, Yilin Sun +4
Unmanned aerial vehicle (UAV) multimodal perception integrates visible (RGB), infrared (IR), synthetic aperture radar (SAR), and depth sensors for scene understanding under diverse…