3 papers
cs.CV2025
RS-MoE: A Vision-Language Model with Mixture of Experts for Remote Sensing Image Captioning and Visual Question Answering
Hui Lin, Danfeng Hong, Shuhang Ge +4
Remote Sensing Image Captioning (RSIC) presents unique challenges and plays a critical role in applications. Traditional RSIC methods often struggle to produce rich and diverse des…
cs.CV2024
LG-VQ: Language-Guided Codebook Learning
Guotao Liang, Baoquan Zhang, Yaowei Wang +6
Vector quantization (VQ) is a key technique in high-resolution and high-fidelity image synthesis, which aims to learn a codebook to encode an image with a sequence of discrete code…
cs.CV2024
MCSDNet: Mesoscale Convective System Detection Network via Multi-scale Spatiotemporal Information
Jiajun Liang, Baoquan Zhang, Yunming Ye +3
The accurate detection of Mesoscale Convective Systems (MCS) is crucial for meteorological monitoring due to their potential to cause significant destruction through severe weather…