2 citations · 2 across the 4 of their papers we have counts for
4 papers
Multi-criteria Token Fusion with One-step-ahead Attention for Efficient Vision Transformers
Sanghyeok Lee, Joonmyung Choi, Hyunwoo J. Kim
Vision Transformer (ViT) has emerged as a prominent backbone for computer vision. For more efficient ViTs, recent works lessen the quadratic cost of the self-attention layer by pru…
vid-TLDR: Training Free Token merging for Light-weight Video Transformer
Joonmyung Choi, Sanghyeok Lee, Jaewon Chu +2
Video Transformers have become the prevalent solution for various video downstream tasks with superior expressive power and flexibility. However, these video transformers suffer fr…
Concept Bottleneck with Visual Concept Filtering for Explainable Medical Image Classification
Injae Kim, Jongha Kim, Joonmyung Choi +1
Interpretability is a crucial factor in building reliable models for various medical applications. Concept Bottleneck Models (CBMs) enable interpretable image classification by uti…
MELTR: Meta Loss Transformer for Learning to Fine-tune Video Foundation Models
Dohwan Ko, Joonmyung Choi, Hyeong Kyu Choi +3
Foundation models have shown outstanding performance and generalization capabilities across domains. Since most studies on foundation models mainly focus on the pretraining phase,…