11 citations · 11 across the 5 of their papers we have counts for
7 papers · 1 filter
Emergence of Text Readability in Vision Language Models
Jaeyoo Park, Sanghyuk Chun, Wonjae Kim +2
We investigate how the ability to recognize textual content within images emerges during the training of Vision-Language Models (VLMs). Our analysis reveals a critical phenomenon:…
Hierarchical Visual Feature Aggregation for OCR-Free Document Understanding
Jaeyoo Park, Jin Young Choi, Jeonghyung Park +1
We present a novel OCR-free document understanding framework based on pretrained Multimodal Large Language Models (MLLMs). Our approach employs multi-scale visual features to effec…
Multi-Modal Representation Learning with Text-Driven Soft Masks
Jaeyoo Park, Bohyung Han
We propose a visual-linguistic representation learning approach within a self-supervised learning framework by introducing a new operation, loss, and data augmentation strategy. Fi…
Cross-Class Feature Augmentation for Class Incremental Learning
Taehoon Kim, Jaeyoo Park, Bohyung Han
We propose a novel class incremental learning approach by incorporating a feature augmentation technique motivated by adversarial attacks. We employ a classifier learned in the pas…
Class-Incremental Learning for Action Recognition in Videos
Jaeyoo Park, Minsoo Kang, Bohyung Han
We tackle catastrophic forgetting problem in the context of class-incremental learning for video recognition, which has not been explored actively despite the popularity of continu…
Learning to Adapt to Unseen Abnormal Activities under Weak Supervision
Jaeyoo Park, Junha Kim, Bohyung Han
We present a meta-learning framework for weakly supervised anomaly detection in videos, where the detector learns to adapt to unseen types of abnormal activities effectively when o…