3 papers
cs.CV2025
PoseLLM: Enhancing Language-Guided Human Pose Estimation with MLP Alignment
Dewen Zhang, Tahir Hussain, Wangpeng An +1
Human pose estimation traditionally relies on architectures that encode keypoint priors, limiting their generalization to novel poses or unseen keypoints. Recent language-guided ap…
cs.CV2025
LLaVA-Pose: Enhancing Human Pose and Action Understanding via Keypoint-Integrated Instruction Tuning
Dewen Zhang, Tahir Hussain, Wangpeng An +1
Current vision-language models (VLMs) are well-adapted for general visual understanding tasks. However, they perform inadequately when handling complex visual tasks related to huma…
cs.CV2025
Keypoint-Integrated Instruction-Following Data Generation for Enhanced Human Pose and Action Understanding in Multimodal Models
Dewen Zhang, Wangpeng An, Hayaru Shouno
Current vision-language multimodal models are well-adapted for general visual understanding tasks. However, they perform inadequately when handling complex visual tasks related to…