11 papers
HyperVAttention: Efficient Sparse Attention with Spatio-Temporal Clustering for Video Diffusion
Dongyeun Lee, Amir Zandieh, Vahab Mirrokni +2
Video Diffusion Transformers (VDiTs) have demonstrated significant capabilities in high-fidelity video generation. However, their ability to produce long-duration videos is fundame…
Learning Neural Deformation Representation for 4D Dynamic Shape Generation
Gyojin Han, Jiwan Hur, Jaehyun Choi +1
Recent developments in 3D shape representation opened new possibilities for generating detailed 3D shapes. Despite these advances, there are few studies dealing with the generation…
B4DL: A Benchmark for 4D LiDAR LLM in Spatio-Temporal Understanding
Changho Choi, Youngwoo Shin, Gyojin Han +2
Understanding dynamic outdoor environments requires capturing complex object interactions and their evolution over time. LiDAR-based 4D point clouds provide precise spatial geometr…
PRISM: Video Dataset Condensation with Progressive Refinement and Insertion for Sparse Motion
Jaehyun Choi, Jiwan Hur, Gyojin Han +2
Video dataset condensation aims to reduce the immense computational cost of video processing. However, it faces a fundamental challenge regarding the inseparable interdependence be…
Inlier-Centric Post-Training Quantization for Object Detection Models
Minsu Kim, Dongyeun Lee, Jaemyung Yu +3
Object detection is pivotal in computer vision, yet its immense computational demands make deployment slow and power-hungry, motivating quantization. However, task-irrelevant morph…
Frequency-Aware Token Reduction for Efficient Vision Transformer
Dong-Jae Lee, Jiwan Hur, Jaehyun Choi +2
Vision Transformers have demonstrated exceptional performance across various computer vision tasks, yet their quadratic computational complexity concerning token length remains a s…