2 papers
cs.CV2026
PulseMind: A Multi-Modal Medical Model for Real-World Clinical Diagnosis
Jiao Xu, Junwei Liu, Jiangwei Lao +9
Recent advances in medical multi-modal models focus on specialized image analysis like dermatology, pathology, or radiology. However, they do not fully capture the complexity of re…
cs.CV2025
When Large Vision-Language Model Meets Large Remote Sensing Imagery: Coarse-to-Fine Text-Guided Token Pruning
Junwei Luo, Yingying Zhang, Xue Yang +5
Efficient vision-language understanding of large Remote Sensing Images (RSIs) is meaningful but challenging. Current Large Vision-Language Models (LVLMs) typically employ limited p…