Fusing Domain-Specific Content from Large Language Models into Knowledge Graphs for Enhanced Zero Shot Object State Classification
arXiv:2403.12151 · doi:10.1609/aaaiss.v3i1.31190
Abstract
Domain-specific knowledge can significantly contribute to addressing a wide variety of vision tasks. However, the generation of such knowledge entails considerable human labor and time costs. This study investigates the potential of Large Language Models (LLMs) in generating and providing domain-specific information through semantic embeddings. To achieve this, an LLM is integrated into a pipeline that utilizes Knowledge Graphs and pre-trained semantic vectors in the context of the Vision-based Zero-shot Object State Classification task. We thoroughly examine the behavior of the LLM through an extensive ablation study. Our findings reveal that the integration of LLM-based embeddings, in combination with general-purpose pre-trained embeddings, leads to substantial performance improvements. Drawing insights from this ablation study, we conduct a comparative analysis against competing models, thereby highlighting the state-of-the-art performance achieved by the proposed approach.
Accepted at the AAAI-MAKE 2024
References in corpus (8)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
- CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment
- How to Generate a Good Word Embedding?
- Learning to Compose Soft Prompts for Compositional Zero-Shot Learning
- Endowing Language Models with Multimodal Knowledge Graph Representations