3 papers
cs.CV2026
StructCore: Structure-Aware Image-Level Scoring for Training-Free Unsupervised Anomaly Detection
Joongwon Chae, Lihui Luo, Yang Liu +8
Max pooling is the de facto standard for converting anomaly score maps into image-level decisions in memory-bank-based unsupervised anomaly detection (UAD). However, because it rel…
cs.CV2024
Grid-augmented vision: A simple yet effective approach for enhanced spatial understanding in multi-modal agents
Joongwon Chae, Zhenyu Wang, Lian Zhang +2
Recent advances in multimodal models have demonstrated impressive capabilities in object recognition and scene understanding. However, these models often struggle with precise spat…
cs.CV2024
SJTU:Spatial judgments in multimodal models towards unified segmentation through coordinate detection
Joongwon Chae, Zhenyu Wang, Peiwu Qin
Despite significant advances in vision-language understanding, implementing image segmentation within multimodal architectures remains a fundamental challenge in modern artificial…