IFShip: Interpretable Fine-grained Ship Classification with Domain Knowledge-Enhanced Vision-Language Models
arXiv:2408.06631 · doi:10.1016/j.patcog.2025.111672
Abstract
End-to-end interpretation currently dominates the remote sensing fine-grained ship classification (RS-FGSC) task. However, the inference process remains uninterpretable, leading to criticisms of these models as "black box" systems. To address this issue, we propose a domain knowledge-enhanced Chain-of-Thought (CoT) prompt generation mechanism, which is used to semi-automatically construct a task-specific instruction-following dataset, TITANIC-FGS. By training on TITANIC-FGS, we adapt general-domain vision-language models (VLMs) to the FGSC task, resulting in a model named IFShip. Building upon IFShip, we develop an FGSC visual chatbot that redefines the FGSC problem as a step-by-step reasoning task and conveys the reasoning process in natural language. Experimental results show that IFShip outperforms state-of-the-art FGSC algorithms in both interpretability and classification accuracy. Furthermore, compared to VLMs such as LLaVA and MiniGPT-4, IFShip demonstrates superior performance on the FGSC task. It provides an accurate chain of reasoning when fine-grained ship types are recognizable to the human eye and offers interpretable explanations when they are not. Our dataset is publicly available at: https://github.com/lostwolves/IFShip.
References in corpus (11)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Visual Instruction Tuning
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
- LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day
- Boosting Few-shot Fine-grained Recognition with Background Suppression and Foreground Alignment
- RSGPT: A Remote Sensing Vision Language Model and Benchmark
- Learning Contrastive Self-Distillation for Ultra-Fine-Grained Visual Categorization Targeting Limited Samples
- MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
- Efficient Multimodal Learning from Data-centric Perspective