3 papers
cs.CV2024
See It All: Contextualized Late Aggregation for 3D Dense Captioning
Minjung Kim, Hyung Suk Lim, Seung Hwan Kim +3
3D dense captioning is a task to localize objects in a 3D scene and generate descriptive sentences for each object. Recent approaches in 3D dense captioning have adopted transforme…
cs.CV2023
Misalign, Contrast then Distill: Rethinking Misalignments in Language-Image Pretraining
Bumsoo Kim, Yeonsik Jo, Jinhyung Kim +1
Contrastive Language-Image Pretraining has emerged as a prominent approach for training vision and text encoders with uncurated image-text pairs from the web. To enhance data-effic…
cs.CV2023
ReConPatch : Contrastive Patch Representation Learning for Industrial Anomaly Detection
Jeeho Hyun, Sangyun Kim, Giyoung Jeon +3
Anomaly detection is crucial to the advanced identification of product defects such as incorrect parts, misaligned components, and damages in industrial manufacturing. Due to the r…