9 papers
Cross-Modal Retrieval for Motion and Text via DropTriple Loss
Sheng Yan, Yang Liu, Haoqiang Wang +3
Cross-modal retrieval of image-text and video-text is a prominent research area in computer vision and natural language processing. However, there has been insufficient attention g…
NanoHTNet: Nano Human Topology Network for Efficient 3D Human Pose Estimation
Jialun Cai, Mengyuan Liu, Hong Liu +2
The widespread application of 3D human pose estimation (HPE) is limited by resource-constrained edge devices, requiring more efficient models. A key approach to enhancing efficienc…
HOT: Hierarchical Hourglass Tokenizer for Efficient Video Pose Transformers
Wenhao Li, Mengyuan Liu, Hong Liu +3
Transformers have been successfully applied in the field of video-based 3D human pose estimation. However, the high computational costs of these video pose transformers (VPTs) make…
Masked Clustering Prediction for Unsupervised Point Cloud Pre-training
Bin Ren, Xiaoshui Huang, Mengyuan Liu +4
Vision transformers (ViTs) have recently been widely applied to 3D point cloud understanding, with masked autoencoding as the predominant pre-training paradigm. However, the challe…
Uncertainty-Aware Testing-Time Optimization for 3D Human Pose Estimation
Ti Wang, Mengyuan Liu, Hong Liu +5
Although data-driven methods have achieved success in 3D human pose estimation, they often suffer from domain gaps and exhibit limited generalization. In contrast, optimization-bas…
Refine Medical Diagnosis Using Generation Augmented Retrieval and Clinical Practice Guidelines
Wenhao Li, Hongkuan Zhang, Hongwei Zhang +5
Current medical language models, adapted from large language models (LLMs), typically predict ICD code-based diagnosis from electronic health records (EHRs) because these labels ar…