3 citations · 7 across the 7 of their papers we have counts for
10 papers · 1 filter
RemoteSAM: Towards Segment Anything for Earth Observation
Liang Yao, Fan Liu, Delong Chen +6
We aim to develop a robust yet flexible visual foundation model for Earth observation. It should possess strong capabilities in recognizing and localizing diverse visual targets wh…
Prompting DirectSAM for Semantic Contour Extraction in Remote Sensing Images
Shiyu Miao, Delong Chen, Fan Liu +4
The Direct Segment Anything Model (DirectSAM) excels in class-agnostic contour extraction. In this paper, we explore its use by applying it to optical remote sensing imagery, where…
Making Large Vision Language Models to be Good Few-shot Learners
Fan Liu, Wenwen Cai, Jian Huo +3
Few-shot classification (FSC) is a fundamental yet challenging task in computer vision that involves recognizing novel classes from limited data. While previous methods have focuse…
What Makes for Good Image Captions?
Delong Chen, Samuel Cahyawijaya, Etsuko Ishii +3
This paper establishes a formal information-theoretic framework for image captioning, conceptualizing captions as compressed linguistic representations that selectively encode sema…
Subobject-level Image Tokenization
Delong Chen, Samuel Cahyawijaya, Jianfeng Liu +2
Patch-based image tokenization ignores the morphology of the visual world, limiting effective and efficient learning of image understanding. Inspired by subword tokenization, we in…
Few-shot Adaptation of Multi-modal Foundation Models: A Survey
Fan Liu, Tianshu Zhang, Wenwen Dai +3
Multi-modal (vision-language) models, such as CLIP, are replacing traditional supervised pre-training models (e.g., ImageNet-based pre-training) as the new generation of visual fou…