activity
20212025
most citedFew-shot Classification via Ensemble Learning with Multi-Order Statistics

3 citations · 7 across the 7 of their papers we have counts for

collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV2025

RemoteSAM: Towards Segment Anything for Earth Observation

Liang Yao, Fan Liu, Delong Chen +6

We aim to develop a robust yet flexible visual foundation model for Earth observation. It should possess strong capabilities in recognizing and localizing diverse visual targets wh…

cs.CV2024

Prompting DirectSAM for Semantic Contour Extraction in Remote Sensing Images

Shiyu Miao, Delong Chen, Fan Liu +4

The Direct Segment Anything Model (DirectSAM) excels in class-agnostic contour extraction. In this paper, we explore its use by applying it to optical remote sensing imagery, where…

cs.CV2024

Making Large Vision Language Models to be Good Few-shot Learners

Fan Liu, Wenwen Cai, Jian Huo +3

Few-shot classification (FSC) is a fundamental yet challenging task in computer vision that involves recognizing novel classes from limited data. While previous methods have focuse…

cs.CV2024

What Makes for Good Image Captions?

Delong Chen, Samuel Cahyawijaya, Etsuko Ishii +3

This paper establishes a formal information-theoretic framework for image captioning, conceptualizing captions as compressed linguistic representations that selectively encode sema…

cs.CV2024

Subobject-level Image Tokenization

Delong Chen, Samuel Cahyawijaya, Jianfeng Liu +2

Patch-based image tokenization ignores the morphology of the visual world, limiting effective and efficient learning of image understanding. Inspired by subword tokenization, we in…

cs.CV2024

Few-shot Adaptation of Multi-modal Foundation Models: A Survey

Fan Liu, Tianshu Zhang, Wenwen Dai +3

Multi-modal (vision-language) models, such as CLIP, are replacing traditional supervised pre-training models (e.g., ImageNet-based pre-training) as the new generation of visual fou…