activity
20242026
collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2025

Towards Efficient Vision State Space Models via Token Merging

Jinyoung Park, Minseok Son, Changick Kim

State Space Models (SSMs) have emerged as powerful architectures in computer vision, yet improving their computational efficiency remains crucial for practical and scalable deploym…

cs.CV2025

Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models

Sangmin Woo, Donguk Kim, Jaehyuk Jang +2

Large Vision Language Models (LVLMs) demonstrate strong capabilities in visual understanding and description, yet often suffer from hallucinations, attributing incorrect or mislead…

cs.CV2024

Difficulty-aware Balancing Margin Loss for Long-tailed Recognition

Minseok Son, Inyong Koo, Jinyoung Park +1

When trained with severely imbalanced data, deep neural networks often struggle to accurately recognize classes with only a few samples. Previous studies in long-tailed recognition…

cs.CV2024

RITUAL: Random Image Transformations as a Universal Anti-hallucination Lever in Large Vision Language Models

Sangmin Woo, Jaehyuk Jang, Donguk Kim +2

Recent advancements in Large Vision Language Models (LVLMs) have revolutionized how machines understand and generate textual responses based on visual inputs, yet they often produc…

cs.CV2024

Diffusion Model Patching via Mixture-of-Prompts

Seokil Ham, Sangmin Woo, Jin-Young Kim +3

We present Diffusion Model Patching (DMP), a simple method to boost the performance of pre-trained diffusion models that have already reached convergence, with a negligible increas…

cs.CV2024

Flow-Assisted Motion Learning Network for Weakly-Supervised Group Activity Recognition

Muhammad Adi Nugroho, Sangmin Woo, Sumin Lee +4

Weakly-Supervised Group Activity Recognition (WSGAR) aims to understand the activity performed together by a group of individuals with the video-level label and without actor-level…