2 papers
cs.CV2024
ImagePiece: Content-aware Re-tokenization for Efficient Image Recognition
Seungdong Yoa, Seungjun Lee, Hyeseung Cho +2
Vision Transformers (ViTs) have achieved remarkable success in various computer vision tasks. However, ViTs have a huge computational cost due to their inherent reliance on multi-h…
cs.CV2024
See It All: Contextualized Late Aggregation for 3D Dense Captioning
Minjung Kim, Hyung Suk Lim, Seung Hwan Kim +3
3D dense captioning is a task to localize objects in a 3D scene and generate descriptive sentences for each object. Recent approaches in 3D dense captioning have adopted transforme…