6 papers
What Makes for Good Image Captions?
Delong Chen, Samuel Cahyawijaya, Etsuko Ishii +3
This paper establishes a formal information-theoretic framework for image captioning, conceptualizing captions as compressed linguistic representations that selectively encode sema…
RemoteSAM: Towards Segment Anything for Earth Observation
Liang Yao, Fan Liu, Delong Chen +6
We aim to develop a robust yet flexible visual foundation model for Earth observation. It should possess strong capabilities in recognizing and localizing diverse visual targets wh…
Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions
Yijun Shen, Delong Chen, Fan Liu +4
While densely annotated image captions significantly facilitate the learning of robust vision-language alignment, methodologies for systematically optimizing human annotation effor…
High-Dimension Human Value Representation in Large Language Models
Samuel Cahyawijaya, Delong Chen, Yejin Bang +5
The widespread application of LLMs across various tasks and fields has necessitated the alignment of these models with human values and preferences. Given various approaches of hum…
Subobject-level Image Tokenization
Delong Chen, Samuel Cahyawijaya, Jianfeng Liu +2
Patch-based image tokenization ignores the morphology of the visual world, limiting effective and efficient learning of image understanding. Inspired by subword tokenization, we in…
Prompting DirectSAM for Semantic Contour Extraction in Remote Sensing Images
Shiyu Miao, Delong Chen, Fan Liu +4
The Direct Segment Anything Model (DirectSAM) excels in class-agnostic contour extraction. In this paper, we explore its use by applying it to optical remote sensing imagery, where…