7 citations · 17 across the 6 of their papers we have counts for
8 papers
Focal and Global Spatial-Temporal Transformer for Skeleton-based Action Recognition
Zhimin Gao, Peitao Wang, Pei Lv +5
Despite great progress achieved by transformer in various vision tasks, it is still underexplored for skeleton-based action recognition with only a few attempts. Besides, these met…
FT-HID: A Large Scale RGB-D Dataset for First and Third Person Human Interaction Analysis
Zihui Guo, Yonghong Hou, Pichao Wang +3
Analysis of human interaction is one important research topic of human motion analysis. It has been studied either using first person vision (FPV) or third person vision (TPV). How…
Trear: Transformer-based RGB-D Egocentric Action Recognition
Xiangyu Li, Yonghong Hou, Pichao Wang +3
In this paper, we propose a \textbf{Tr}ansformer-based RGB-D \textbf{e}gocentric \textbf{a}ction \textbf{r}ecognition framework, called Trear. It consists of two modules, inter-fra…
Transformer Guided Geometry Model for Flow-Based Unsupervised Visual Odometry
Xiangyu Li, Yonghong Hou, Pichao Wang +3
Existing unsupervised visual odometry (VO) methods either match pairwise images or integrate the temporal information using recurrent neural networks over a long sequence of images…
MDSSD: Multi-scale Deconvolutional Single Shot Detector for Small Objects
Lisha Cui, Rui Ma, Pei Lv +4
For most of the object detectors based on multi-scale feature maps, the shallow layers are rich in fine spatial information and thus mainly responsible for small object detection.…
Depth Pooling Based Large-scale 3D Action Recognition with Convolutional Neural Networks
Pichao Wang, Wanqing Li, Zhimin Gao +2
This paper proposes three simple, compact yet effective representations of depth sequences, referred to respectively as Dynamic Depth Images (DDI), Dynamic Depth Normal Images (DDN…