1 paper
Kexin Ma, Jing Xiao, Bowen Xing +2
RS-LVLMs have advanced multimodal understanding of Earth observation imagery, yet their performance is fundamentally constrained by high-resolution processing, as visual token coun…