activity
20152022
most citedTrear: Transformer-based RGB-D Egocentric Action Recognition

7 citations · 17 across the 6 of their papers we have counts for

collaborators

8 papers

cs.CV20225 cited

Focal and Global Spatial-Temporal Transformer for Skeleton-based Action Recognition

Zhimin Gao, Peitao Wang, Pei Lv +5

Despite great progress achieved by transformer in various vision tasks, it is still underexplored for skeleton-based action recognition with only a few attempts. Besides, these met…

cs.CV2022

FT-HID: A Large Scale RGB-D Dataset for First and Third Person Human Interaction Analysis

Zihui Guo, Yonghong Hou, Pichao Wang +3

Analysis of human interaction is one important research topic of human motion analysis. It has been studied either using first person vision (FPV) or third person vision (TPV). How…

cs.CV20217 cited

Trear: Transformer-based RGB-D Egocentric Action Recognition

Xiangyu Li, Yonghong Hou, Pichao Wang +3

In this paper, we propose a \textbf{Tr}ansformer-based RGB-D \textbf{e}gocentric \textbf{a}ction \textbf{r}ecognition framework, called Trear. It consists of two modules, inter-fra…

cs.CV2020

Transformer Guided Geometry Model for Flow-Based Unsupervised Visual Odometry

Xiangyu Li, Yonghong Hou, Pichao Wang +3

Existing unsupervised visual odometry (VO) methods either match pairwise images or integrate the temporal information using recurrent neural networks over a long sequence of images…

cs.CV2018

MDSSD: Multi-scale Deconvolutional Single Shot Detector for Small Objects

Lisha Cui, Rui Ma, Pei Lv +4

For most of the object detectors based on multi-scale feature maps, the shallow layers are rich in fine spatial information and thus mainly responsible for small object detection.…

cs.CV2018

Depth Pooling Based Large-scale 3D Action Recognition with Convolutional Neural Networks

Pichao Wang, Wanqing Li, Zhimin Gao +2

This paper proposes three simple, compact yet effective representations of depth sequences, referred to respectively as Dynamic Depth Images (DDI), Dynamic Depth Normal Images (DDN…