11 citations · 11 across the 4 of their papers we have counts for
4 papers
DualFocus: Integrating Macro and Micro Perspectives in Multi-modal Large Language Models
Yuhang Cao, Pan Zhang, Xiaoyi Dong +2
We present DualFocus, a novel framework for integrating macro and micro perspectives within multi-modal large language models (MLLMs) to enhance vision-language task performance. C…
InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
Xiaoyi Dong, Pan Zhang, Yuhang Zang +20
We introduce InternLM-XComposer2, a cutting-edge vision-language model excelling in free-form text-image composition and comprehension. This model goes beyond conventional vision-l…
HGDNet: A Height-Hierarchy Guided Dual-Decoder Network for Single View Building Extraction and Height Estimation
Chaoran Lu, Ningning Cao, Pan Zhang +7
Unifying the correlative single-view satellite image building extraction and height estimation tasks indicates a promising way to share representations and acquire generalist model…
Fine-grained building roof instance segmentation based on domain adapted pretraining and composite dual-backbone
Guozhang Liu, Baochai Peng, Ting Liu +7
The diversity of building architecture styles of global cities situated on various landforms, the degraded optical imagery affected by clouds and shadows, and the significant inter…