39 citations · 41 across the 3 of their papers we have counts for
3 papers · 1 filter
Structured Redundancy Modeling for Efficient Visual Token Pruning in High-Resolution MLLMs
Jouwon Song, Woohyeong Kim, Kyeongbo Kong
Recent high-resolution Multimodal Large Language Models (MLLMs) generate thousands of visual tokens per input, leading to a visual token explosion that introduces severe latency bo…
AgilePruner: An Empirical Study of Attention and Diversity for Adaptive Visual Token Pruning in Large Vision-Language Models
Changwoo Baek, Jouwon Song, Sohyeon Kim +1
Large Vision-Language Models (LVLMs) have adopted visual token pruning strategies to mitigate substantial computational overhead incurred by extensive visual token sequences. While…
Attention Map-guided Two-stage Anomaly Detection using Hard Augmentation
Jou Won Song, Kyeongbo Kong, Ye In Park +1
Anomaly detection is a task that recognizes whether an input sample is included in the distribution of a target normal class or an anomaly class. Conventional generative adversaria…