activity
20242026
collaborators

9 papers

cs.CV2026

Cross-Modal Retrieval for Motion and Text via DropTriple Loss

Sheng Yan, Yang Liu, Haoqiang Wang +3

Cross-modal retrieval of image-text and video-text is a prominent research area in computer vision and natural language processing. However, there has been insufficient attention g…

cs.CV2025

NanoHTNet: Nano Human Topology Network for Efficient 3D Human Pose Estimation

Jialun Cai, Mengyuan Liu, Hong Liu +2

The widespread application of 3D human pose estimation (HPE) is limited by resource-constrained edge devices, requiring more efficient models. A key approach to enhancing efficienc…

cs.CV2025

HOT: Hierarchical Hourglass Tokenizer for Efficient Video Pose Transformers

Wenhao Li, Mengyuan Liu, Hong Liu +3

Transformers have been successfully applied in the field of video-based 3D human pose estimation. However, the high computational costs of these video pose transformers (VPTs) make…

cs.CV2025

Masked Clustering Prediction for Unsupervised Point Cloud Pre-training

Bin Ren, Xiaoshui Huang, Mengyuan Liu +4

Vision transformers (ViTs) have recently been widely applied to 3D point cloud understanding, with masked autoencoding as the predominant pre-training paradigm. However, the challe…

cs.CV2025

Uncertainty-Aware Testing-Time Optimization for 3D Human Pose Estimation

Ti Wang, Mengyuan Liu, Hong Liu +5

Although data-driven methods have achieved success in 3D human pose estimation, they often suffer from domain gaps and exhibit limited generalization. In contrast, optimization-bas…

cs.CL2025

Refine Medical Diagnosis Using Generation Augmented Retrieval and Clinical Practice Guidelines

Wenhao Li, Hongkuan Zhang, Hongwei Zhang +5

Current medical language models, adapted from large language models (LLMs), typically predict ICD code-based diagnosis from electronic health records (EHRs) because these labels ar…