2 papers
cs.CV2026
Efficient Reasoning via Thought Compression for Language Segmentation
Qing Zhou, Shiyu Zhang, Yuyu Jia +4
Chain-of-thought (CoT) reasoning has significantly improved the performance of large multimodal models in language-guided segmentation, yet its prohibitive computational cost, stem…
cs.CV2025
A Benchmark for Multi-Lingual Vision-Language Learning in Remote Sensing Image Captioning
Qing Zhou, Tao Yang, Junyu Gao +3
Remote Sensing Image Captioning (RSIC) is a cross-modal field bridging vision and language, aimed at automatically generating natural language descriptions of features and scenes i…