10 papers
Context-structured Video Anomaly Detection with Large Vision-Language Models
Dongjun Kim, Changjae Oh, Andrea Cavallaro +1
Training video anomaly detectors is challenging due to the difficulty and cost of annotating diverse and rare abnormal events. Although recent large vision-language models enable t…
FlowOVD: Learning Generative Latent Flows for Zero-shot Open-vocabulary Detection
Yao Wei, Andrea Cavallaro, Changjae Oh
Open-vocabulary object detection (OVD) has achieved remarkable progress through large-scale vision-language pre-training. Existing methods, however, typically formulate OVD as a di…
The Detector Teaches Itself: Lightweight Self-Supervised Adaptation for Open-Vocabulary Object Detection
Yazhe Wan, Changjae Oh
Open-vocabulary object detection aims to recognize objects from an open set of categories, which leverages vision-language models (VLMs) pre-trained on large-scale image-text data.…
Chain-of-Caption: Training-free improvement of multimodal large language model on referring expression comprehension
Yik Lung Pang, Changjae Oh
Given a textual description, the task of referring expression comprehension (REC) involves the localisation of the referred object in an image. Multimodal large language models (ML…
AutoMV: An Automatic Multi-Agent System for Music Video Generation
Xiaoxuan Tang, Xinping Lei, Chaoran Zhu +10
Music-to-Video (M2V) generation for full-length songs faces significant challenges. Existing methods produce short, disjointed clips, failing to align visuals with musical structur…
LaVA-Man: Learning Visual Action Representations for Robot Manipulation
Chaoran Zhu, Hengyi Wang, Yik Lung Pang +1
Visual-textual understanding is essential for language-guided robot manipulation. Recent works leverage pre-trained vision-language models to measure the similarity between encoded…