6 papers
Optimization of Prompt Learning via Multi-Knowledge Representation for Vision-Language Models
Enming Zhang, Bingke Zhu, Yingying Chen +3
Vision-Language Models (VLMs), such as CLIP, play a foundational role in various cross-modal applications. To fully leverage VLMs' potential in adapting to downstream tasks, contex…
Friend or Foe? Harnessing Controllable Overfitting for Anomaly Detection
Long Qian, Bingke Zhu, Yingying Chen +2
Overfitting has traditionally been viewed as detrimental to anomaly detection, where excessive generalization often limits models' sensitivity to subtle anomalies. Our work challen…
UniVAD: A Training-free Unified Model for Few-shot Visual Anomaly Detection
Zhaopeng Gu, Bingke Zhu, Guibo Zhu +3
Visual Anomaly Detection (VAD) aims to identify abnormal samples in images that deviate from normal patterns, covering multiple domains, including industrial, logical, and medical…
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing
Langyu Wang, Bingke Zhu, Yingying Chen +1
Audio-visual video parsing focuses on classifying videos through weak labels while identifying events as either visible, audible, or both, alongside their respective temporal bound…
MROVSeg: Breaking the Resolution Curse of Vision-Language Models in Open-Vocabulary Image Segmentation
Yuanbing Zhu, Bingke Zhu, Yingying Chen +3
Pretrained vision-language models (VLMs), \eg CLIP, are increasingly used to bridge the gap between open- and close-vocabulary recognition in open-vocabulary image segmentation. As…
The BRAVO Semantic Segmentation Challenge Results in UNCV2024
Tuan-Hung Vu, Eduardo Valle, Andrei Bursuc +16
We propose the unified BRAVO challenge to benchmark the reliability of semantic segmentation models under realistic perturbations and unknown out-of-distribution (OOD) scenarios. W…