2 papers
cs.CV2025
EarthGPT-X: A Spatial MLLM for Multi-level Multi-Source Remote Sensing Imagery Understanding with Visual Prompting
Wei Zhang, Miaoxin Cai, Yaqian Ning +6
Recent advances in natural-domain multi-modal large language models (MLLMs) have demonstrated effective spatial reasoning through visual and textual prompting. However, their direc…
cs.CV2025
UniDet-D: A Unified Dynamic Spectral Attention Model for Object Detection under Adverse Weathers
Wei Zhang, Yuantao Wang, Haowei Yang +3
Real-world object detection is a challenging task where the captured images/videos often suffer from complex degradations due to various adverse weather conditions such as rain, fo…