6 papers
Messages, Not Tokens: Grounded Coresets for Faithful VLM Compression
Long Qian, Jiaqi Wei, Bingke Zhu +2
Modern vision language models (VLMs) turn high-resolution images into long sequences of visual tokens. Every token traverses the language decoder and persists in its prompt KV cach…
SmoothGuard: Defending Multimodal Large Language Models with Noise Perturbation and Clustering Aggregation
Guangzhi Su, Shuchang Huang, Yutong Ke +3
Multimodal large language models (MLLMs) have achieved impressive performance across diverse tasks by jointly reasoning over textual and visual inputs. Despite their success, these…
Quality-Aware Language-Conditioned Local Auto-Regressive Anomaly Synthesis and Detection
Long Qian, Bingke Zhu, Yingying Chen +2
Despite substantial progress in anomaly synthesis methods, existing diffusion-based and coarse inpainting pipelines commonly suffer from structural deficiencies such as micro-struc…
MathPhys-Guided Coarse-to-Fine Anomaly Synthesis with SQE-Driven Bi-Level Optimization for Anomaly Detection
Long Qian, Bingke Zhu, Yingying Chen +2
Currently, industrial anomaly detection suffers from two bottlenecks: (i) the rarity of real-world defect images and (ii) the opacity of sample quality when synthetic data are used…
Friend or Foe? Harnessing Controllable Overfitting for Anomaly Detection
Long Qian, Bingke Zhu, Yingying Chen +2
Overfitting has traditionally been viewed as detrimental to anomaly detection, where excessive generalization often limits models' sensitivity to subtle anomalies. Our work challen…
The BRAVO Semantic Segmentation Challenge Results in UNCV2024
Tuan-Hung Vu, Eduardo Valle, Andrei Bursuc +16
We propose the unified BRAVO challenge to benchmark the reliability of semantic segmentation models under realistic perturbations and unknown out-of-distribution (OOD) scenarios. W…