collaborators

6 papers

cs.CV2025

What Makes for Good Image Captions?

Delong Chen, Samuel Cahyawijaya, Etsuko Ishii +3

This paper establishes a formal information-theoretic framework for image captioning, conceptualizing captions as compressed linguistic representations that selectively encode sema…

cs.CV2025

RemoteSAM: Towards Segment Anything for Earth Observation

Liang Yao, Fan Liu, Delong Chen +6

We aim to develop a robust yet flexible visual foundation model for Earth observation. It should possess strong capabilities in recognizing and localizing diverse visual targets wh…

cs.CL2025

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions

Yijun Shen, Delong Chen, Fan Liu +4

While densely annotated image captions significantly facilitate the learning of robust vision-language alignment, methodologies for systematically optimizing human annotation effor…

cs.CL2025

High-Dimension Human Value Representation in Large Language Models

Samuel Cahyawijaya, Delong Chen, Yejin Bang +5

The widespread application of LLMs across various tasks and fields has necessitated the alignment of these models with human values and preferences. Given various approaches of hum…

cs.CV2025

Subobject-level Image Tokenization

Delong Chen, Samuel Cahyawijaya, Jianfeng Liu +2

Patch-based image tokenization ignores the morphology of the visual world, limiting effective and efficient learning of image understanding. Inspired by subword tokenization, we in…

cs.CV2024

Prompting DirectSAM for Semantic Contour Extraction in Remote Sensing Images

Shiyu Miao, Delong Chen, Fan Liu +4

The Direct Segment Anything Model (DirectSAM) excels in class-agnostic contour extraction. In this paper, we explore its use by applying it to optical remote sensing imagery, where…