6 papers
HOT: Hierarchical Hourglass Tokenizer for Efficient Video Pose Transformers
Wenhao Li, Mengyuan Liu, Hong Liu +3
Transformers have been successfully applied in the field of video-based 3D human pose estimation. However, the high computational costs of these video pose transformers (VPTs) make…
Masked Clustering Prediction for Unsupervised Point Cloud Pre-training
Bin Ren, Xiaoshui Huang, Mengyuan Liu +4
Vision transformers (ViTs) have recently been widely applied to 3D point cloud understanding, with masked autoencoding as the predominant pre-training paradigm. However, the challe…
Uncertainty-Aware Testing-Time Optimization for 3D Human Pose Estimation
Ti Wang, Mengyuan Liu, Hong Liu +5
Although data-driven methods have achieved success in 3D human pose estimation, they often suffer from domain gaps and exhibit limited generalization. In contrast, optimization-bas…
Fractal-IR: A Unified Framework for Efficient and Scalable Image Restoration
Yawei Li, Bin Ren, Jingyun Liang +5
While vision transformers achieve significant breakthroughs in various image restoration (IR) tasks, it is still challenging to efficiently scale them across multiple types of degr…
Sharing Key Semantics in Transformer Makes Efficient Image Restoration
Bin Ren, Yawei Li, Jingyun Liang +6
Image Restoration (IR), a classic low-level vision task, has witnessed significant advancements through deep models that effectively model global information. Notably, the emergenc…
Hierarchical Information Flow for Generalized Efficient Image Restoration
Yawei Li, Bin Ren, Jingyun Liang +5
While vision transformers show promise in numerous image restoration (IR) tasks, the challenge remains in efficiently generalizing and scaling up a model for multiple IR tasks. To…