6 papers · 1 filter
PulseMind: A Multi-Modal Medical Model for Real-World Clinical Diagnosis
Jiao Xu, Junwei Liu, Jiangwei Lao +9
Recent advances in medical multi-modal models focus on specialized image analysis like dermatology, pathology, or radiology. However, they do not fully capture the complexity of re…
Cross-View Geo-Localization with Street-View and VHR Satellite Imagery in Decentrality Settings
Panwang Xia, Lei Yu, Yi Wan +12
Cross-View Geo-Localization tackles the challenge of image geo-localization in GNSS-denied environments, including disaster response scenarios, urban canyons, and dense forests, by…
HomoMatcher: Dense Feature Matching Results with Semi-Dense Efficiency by Homography Estimation
Xiaolong Wang, Lei Yu, Yingying Zhang +6
Feature matching between image pairs is a fundamental problem in computer vision that drives many applications, such as SLAM. Recently, semi-dense matching approaches have achieved…
POA: Pre-training Once for Models of All Sizes
Yingying Zhang, Xin Guo, Jiangwei Lao +7
Large-scale self-supervised pre-training has paved the way for one foundation model to handle many different vision tasks. Most pre-training methodologies train a single model of a…
SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding
Junwei Luo, Zhen Pang, Yongjun Zhang +8
Remote Sensing Large Multi-Modal Models (RSLMMs) are developing rapidly and showcase significant capabilities in remote sensing imagery (RSI) comprehension. However, due to the lim…
Enhancing DETRs Variants through Improved Content Query and Similar Query Aggregation
Yingying Zhang, Chuangji Shi, Xin Guo +4
The design of the query is crucial for the performance of DETR and its variants. Each query consists of two components: a content part and a positional one. Traditionally, the cont…