28 citations · 88 across the 11 of their papers we have counts for
11 papers
Peeling the Onion: Hierarchical Reduction of Data Redundancy for Efficient Vision Transformer Training
Zhenglun Kong, Haoyu Ma, Geng Yuan +12
Vision transformers (ViTs) have recently obtained success in many applications, but their intensive computation and heavy memory usage at both training and inference time limit the…
Identity-Aware Hand Mesh Estimation and Personalization from RGB Images
Deying Kong, Linguang Zhang, Liangjian Chen +6
Reconstructing 3D hand meshes from monocular RGB images has attracted increasing amount of attention due to its enormous potential applications in the field of AR/VR. Most state-of…
PPT: token-Pruned Pose Transformer for monocular and multi-view human pose estimation
Haoyu Ma, Zhe Wang, Yifei Chen +6
Recently, the vision transformer and its variants have played an increasingly important role in both monocular and multi-view human pose estimation. Considering image patches as to…
Sparsity Winning Twice: Better Robust Generalization from More Efficient Training
Tianlong Chen, Zhenyu Zhang, Pengjun Wang +4
Recent studies demonstrate that deep networks, even robustified by the state-of-the-art adversarial training (AT), still suffer from large robust generalization gaps, in addition t…
VAQF: Fully Automatic Software-Hardware Co-Design Framework for Low-Bit Vision Transformer
Mengshu Sun, Haoyu Ma, Guoliang Kang +5
The transformer architectures with attention mechanisms have obtained success in Nature Language Processing (NLP), and Vision Transformers (ViTs) have recently extended the applica…
AFTer-UNet: Axial Fusion Transformer UNet for Medical Image Segmentation
Xiangyi Yan, Hao Tang, Shanlin Sun +3
Recent advances in transformer-based models have drawn attention to exploring these techniques in medical image segmentation, especially in conjunction with the U-Net model (or its…