activity
20242026
most citedLLaVA-Pose: Enhancing Human Pose and Action Understanding via Keypoint-Integrated Instruction Tuning

1 citations · 1 across the 5 of their papers we have counts for

collaborators

7 papers

physics.med-ph2026

Computed Tomography Reconstruction Algorithm Using Markov Random Field Model

Taiga Shimomiya, Taichi Kusumi, Masayuki Uesugi +4

X-ray computed tomography (CT) reveals the materials' internal structures non-destructively from a tilt series of projected images. Filtered back projection (FBP) is a widely-adopt…

cond-mat.str-el2025

Development of ultra-high efficiency soft X-ray angle-resolved photoemission spectroscopy equipped with deep prior-based denoising method

Kohei Yamagami, Yuichi Yokoyama, Yuta Sumiya +3

Soft X-ray angle resolved photoemission spectroscopy (SX-ARPES) is one of the most powerful spectroscopic techniques to visualize the three-dimensional bulk electronic structure in…

cond-mat.mtrl-sci2025

Deep prior-based denoising for state-of-the-art scientific imaging and metrology

Yuichi Yokoyama, Kohei Yamagami, Yuta Sumiya +2

Deep learning has revolutionized computer vision, yet a major gap persists between complex, data-hungry deep learning models and the practical demands of state-of-the-art scientifi…

cs.CV2025

PoseLLM: Enhancing Language-Guided Human Pose Estimation with MLP Alignment

Dewen Zhang, Tahir Hussain, Wangpeng An +1

Human pose estimation traditionally relies on architectures that encode keypoint priors, limiting their generalization to novel poses or unseen keypoints. Recent language-guided ap…

cs.CV20251 cited

LLaVA-Pose: Enhancing Human Pose and Action Understanding via Keypoint-Integrated Instruction Tuning

Dewen Zhang, Tahir Hussain, Wangpeng An +1

Current vision-language models (VLMs) are well-adapted for general visual understanding tasks. However, they perform inadequately when handling complex visual tasks related to huma…

cs.CV2024

Keypoint-Integrated Instruction-Following Data Generation for Enhanced Human Pose and Action Understanding in Multimodal Models

Dewen Zhang, Wangpeng An, Hayaru Shouno

Current vision-language multimodal models are well-adapted for general visual understanding tasks. However, they perform inadequately when handling complex visual tasks related to…