7 papers · 1 filter
IWP: Token Pruning as Implicit Weight Pruning in Large Vision Language Models
Dong-Jae Lee, Sunghyun Baek, Junmo Kim
Large Vision Language Models show impressive performance across image and video understanding tasks, yet their computational cost grows rapidly with the number of visual tokens. Ex…
Frequency-Aware Token Reduction for Efficient Vision Transformer
Dong-Jae Lee, Jiwan Hur, Jaehyun Choi +2
Vision Transformers have demonstrated exceptional performance across various computer vision tasks, yet their quadratic computational complexity concerning token length remains a s…
B4DL: A Benchmark for 4D LiDAR LLM in Spatio-Temporal Understanding
Changho Choi, Youngwoo Shin, Gyojin Han +2
Understanding dynamic outdoor environments requires capturing complex object interactions and their evolution over time. LiDAR-based 4D point clouds provide precise spatial geometr…
DAM: Domain-Aware Module for Multi-Domain Dataset Condensation
Jaehyun Choi, Gyojin Han, Dong-Jae Lee +2
Dataset Condensation (DC) has emerged as a promising solution to mitigate the computational and storage burdens associated with training deep learning models. However, existing DC…
Self-supervised Transformation Learning for Equivariant Representations
Jaemyung Yu, Jaehyun Choi, Dong-Jae Lee +2
Unsupervised representation learning has significantly advanced various machine learning tasks. In the computer vision domain, state-of-the-art approaches utilize transformations l…
Unlocking the Capabilities of Masked Generative Models for Image Synthesis via Self-Guidance
Jiwan Hur, Dong-Jae Lee, Gyojin Han +3
Masked generative models (MGMs) have shown impressive generative ability while providing an order of magnitude efficient sampling steps compared to continuous diffusion models. How…