paper

Toward Universal Skeleton-Based Action Recognition across Heterogeneous Skeletons and Open Vocabularies

arXiv:2604.17013

Abstract

Skeleton data used for action recognition are acquired from a wide range of sources, including depth sensors, marker-based motion capture systems, and 2D/3D pose estimators. These sources yield skeletons that differ in joint number, skeletal topology, and coordinate dimensionality, making skeleton data inherently heterogeneous. However, previous works overlook the data heterogeneity of skeletons and solely construct models using homogeneous skeletons. Moreover, open-vocabulary action recognition is also essential for real-world applications. To this end, this work studies the challenging problem of heterogeneous skeleton-based action recognition with open vocabularies. We construct a large-scale Heterogeneous Open-Vocabulary (HOV) Skeleton dataset by integrating and refining multiple representative large-scale skeleton-based action datasets. To address universal skeleton-based action recognition, we propose a Transformer-based model that standardizes heterogeneous skeletons into a unified representation, encodes multi-modal skeleton embeddings with a two-stream motion encoder to learn spatio-temporal action representations, and maps them to a semantic space through multi-grained motion-text alignment. The alignment incorporates contrastive learning at three levels: global instance alignment, stream-specific alignment, and fine-grained alignment. Extensive experiments on popular benchmarks with heterogeneous skeleton data demonstrate both the effectiveness and the generalization ability of the proposed method. Code is available at https://github.com/jidongkuang/Universal-Skeleton.

Toward Universal Skeleton-Based Action Recognition across Heterogeneous Skeletons and Open Vocabularies · wovepaper