5 citations · 6 across the 5 of their papers we have counts for
5 papers
Robust Multimodal 3D Object Detection via Modality-Agnostic Decoding and Proximity-based Modality Ensemble
Juhan Cha, Minseok Joo, Jihwan Park +3
Recent advancements in 3D object detection have benefited from multi-modal information from the multi-view cameras and LiDAR sensors. However, the inherent disparities between the…
Multi-criteria Token Fusion with One-step-ahead Attention for Efficient Vision Transformers
Sanghyeok Lee, Joonmyung Choi, Hyunwoo J. Kim
Vision Transformer (ViT) has emerged as a prominent backbone for computer vision. For more efficient ViTs, recent works lessen the quadratic cost of the self-attention layer by pru…
vid-TLDR: Training Free Token merging for Light-weight Video Transformer
Joonmyung Choi, Sanghyeok Lee, Jaewon Chu +2
Video Transformers have become the prevalent solution for various video downstream tasks with superior expressive power and flexibility. However, these video transformers suffer fr…
Read-only Prompt Optimization for Vision-Language Few-shot Learning
Dongjun Lee, Seokwon Song, Jihee Suh +3
In recent years, prompt tuning has proven effective in adapting pre-trained vision-language models to downstream tasks. These methods aim to adapt the pre-trained models by introdu…
Self-positioning Point-based Transformer for Point Cloud Understanding
Jinyoung Park, Sanghyeok Lee, Sihyeon Kim +2
Transformers have shown superior performance on various computer vision tasks with their capabilities to capture long-range dependencies. Despite the success, it is challenging to…