1 paper
Jingpu Yang, Fengxian Ji, Mingxuan Cui +4
UAV video question answering requires separating camera motion from changes in the scene, but RGB-only multimodal models receive no explicit, stable reference for that separation.…