9 papers
IWP: Token Pruning as Implicit Weight Pruning in Large Vision Language Models
Dong-Jae Lee, Sunghyun Baek, Junmo Kim
Large Vision Language Models show impressive performance across image and video understanding tasks, yet their computational cost grows rapidly with the number of visual tokens. Ex…
Continuous Audio Thinking for Large Audio Language Models
Gyojin Han, Dong-Jae Lee, Changho Choi +2
Large audio language models (LALMs) have shown impressive capabilities on diverse audio understanding tasks, ranging from speech transcription to music analysis. However, because L…
B4DL: A Benchmark for 4D LiDAR LLM in Spatio-Temporal Understanding
Changho Choi, Youngwoo Shin, Gyojin Han +2
Understanding dynamic outdoor environments requires capturing complex object interactions and their evolution over time. LiDAR-based 4D point clouds provide precise spatial geometr…
Frequency-Aware Token Reduction for Efficient Vision Transformer
Dong-Jae Lee, Jiwan Hur, Jaehyun Choi +2
Vision Transformers have demonstrated exceptional performance across various computer vision tasks, yet their quadratic computational complexity concerning token length remains a s…
SynAD: Enhancing Real-World End-to-End Autonomous Driving Models through Synthetic Data Integration
Jongsuk Kim, Jaeyoung Lee, Gyojin Han +3
Recent advancements in deep learning and the availability of high-quality real-world driving datasets have propelled end-to-end autonomous driving. Despite this progress, relying s…
DAM: Domain-Aware Module for Multi-Domain Dataset Condensation
Jaehyun Choi, Gyojin Han, Dong-Jae Lee +2
Dataset Condensation (DC) has emerged as a promising solution to mitigate the computational and storage burdens associated with training deep learning models. However, existing DC…