7 papers · 1 filter
Messages, Not Tokens: Grounded Coresets for Faithful VLM Compression
Long Qian, Jiaqi Wei, Bingke Zhu +2
Modern vision language models (VLMs) turn high-resolution images into long sequences of visual tokens. Every token traverses the language decoder and persists in its prompt KV cach…
UniVAD v2: Unified Visual Anomaly Detection via Support-Conditioned Boundary Construction
Zhaopeng Gu, Bingke Zhu, Zhaowen Li +5
Unified visual anomaly detection seeks to train a single detector that can be deployed across categories, domains, and application scenarios. In the few-shot transfer regime, the k…
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing
Langyu Wang, Bingke Zhu, Yingying Chen +3
The weakly-supervised audio-visual video parsing (AVVP) aims to predict all modality-specific events and locate their temporal boundaries. Despite significant progress, due to the…
AnomalyMoE: Towards a Language-free Generalist Model for Unified Visual Anomaly Detection
Zhaopeng Gu, Bingke Zhu, Guibo Zhu +4
Anomaly detection is a critical task across numerous domains and modalities, yet existing methods are often highly specialized, limiting their generalizability. These specialized m…
Quality-Aware Language-Conditioned Local Auto-Regressive Anomaly Synthesis and Detection
Long Qian, Bingke Zhu, Yingying Chen +2
Despite substantial progress in anomaly synthesis methods, existing diffusion-based and coarse inpainting pipelines commonly suffer from structural deficiencies such as micro-struc…
MathPhys-Guided Coarse-to-Fine Anomaly Synthesis with SQE-Driven Bi-Level Optimization for Anomaly Detection
Long Qian, Bingke Zhu, Yingying Chen +2
Currently, industrial anomaly detection suffers from two bottlenecks: (i) the rarity of real-world defect images and (ii) the opacity of sample quality when synthetic data are used…