24 citations · 24 across the 2 of their papers we have counts for
4 papers · 1 filter
Manboformer: Learning Gaussian Representations via Spatial-temporal Attention Mechanism
Ziyue Zhao, Qining Qi, Jianfa Ma
Compared with voxel-based grid prediction, in the field of 3D semantic occupation prediction for autonomous driving, GaussianFormer proposed using 3D Gaussian to describe scenes wi…
AEGIS-Net: Attention-guided Multi-Level Feature Aggregation for Indoor Place Recognition
Yuhang Ming, Jian Ma, Xingrui Yang +3
We present AEGIS-Net, a novel indoor place recognition model that takes in RGB point clouds and generates global place descriptors by aggregating lower-level color, geometry featur…
EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations
Ahmad Darkhalil, Dandan Shan, Bin Zhu +6
We introduce VISOR, a new dataset of pixel annotations and a benchmark suite for segmenting hands and active objects in egocentric video. VISOR annotates videos from EPIC-KITCHENS,…
Hand-Object Interaction Reasoning
Jian Ma, Dima Damen
This paper proposes an interaction reasoning network for modelling spatio-temporal relationships between hands and objects in video. The proposed interaction unit utilises a Transf…