7 papers · 1 filter
Context-structured Video Anomaly Detection with Large Vision-Language Models
Dongjun Kim, Changjae Oh, Andrea Cavallaro +1
Training video anomaly detectors is challenging due to the difficulty and cost of annotating diverse and rare abnormal events. Although recent large vision-language models enable t…
FlowOVD: Learning Generative Latent Flows for Zero-shot Open-vocabulary Detection
Yao Wei, Andrea Cavallaro, Changjae Oh
Open-vocabulary object detection (OVD) has achieved remarkable progress through large-scale vision-language pre-training. Existing methods, however, typically formulate OVD as a di…
The Detector Teaches Itself: Lightweight Self-Supervised Adaptation for Open-Vocabulary Object Detection
Yazhe Wan, Changjae Oh
Open-vocabulary object detection aims to recognize objects from an open set of categories, which leverages vision-language models (VLMs) pre-trained on large-scale image-text data.…
Chain-of-Caption: Training-free improvement of multimodal large language model on referring expression comprehension
Yik Lung Pang, Changjae Oh
Given a textual description, the task of referring expression comprehension (REC) involves the localisation of the referred object in an image. Multimodal large language models (ML…
High-resolution open-vocabulary object 6D pose estimation
Jaime Corsetti, Davide Boscaini, Francesco Giuliari +3
The generalisation to unseen objects in the 6D pose estimation task is very challenging. While Vision-Language Models (VLMs) enable using natural language descriptions to support 6…
Open-vocabulary object 6D pose estimation
Jaime Corsetti, Davide Boscaini, Changjae Oh +2
We introduce the new setting of open-vocabulary object 6D pose estimation, in which a textual prompt is used to specify the object of interest. In contrast to existing approaches,…