1.2k citations
- Nanjing UniversityCN179 papers
- Zhejiang UniversityCN139 papers
- Shanghai Jiao Tong UniversityCN135 papers
- Peking UniversityCN134 papers
- Tsinghua UniversityCN133 papers
- University of Science and Technology of ChinaCN130 papers
- University of Chinese Academy of SciencesCN123 papers
- Fudan UniversityCN119 papers
- Institute of High Energy PhysicsCN117 papers
- Lanzhou UniversityCN117 papers
- Sun Yat-sen UniversityCN117 papers
- Nanjing Normal UniversityCN116 papers
18 papers · 2 filters
Multimodal Skeleton-Based Action Representation Learning via Decomposition and Composition
Hongsong Wang, Heng Fei, Bingxuan Dai +1
Multimodal human action understanding is a significant problem in computer vision, with the central challenge being the effective utilization of the complementarity among diverse m…
Exploring Spatial-Temporal Representation via Star Graph for mmWave Radar-based Human Activity Recognition
Senhao Gao, Junqing Zhang, Luoyu Mei +2
Human activity recognition (HAR) requires extracting accurate spatial-temporal features with human movements. A mmWave radar point cloud-based HAR system suffers from sparsity and…
Gestura: A LVLM-Powered System Bridging Motion and Semantics for Real-Time Free-Form Gesture Understanding
Zhuoming Li, Aitong Liu, Mengxi Jia +5
Free-form gesture understanding is highly appealing for human-computer interaction, as it liberates users from the constraints of predefined gesture categories. However, the sole e…
Improving Generalized Visual Grounding with Instance-aware Joint Learning
Ming Dai, Wenxuan Cheng, Jiang-Jiang Liu +4
Generalized visual grounding tasks, including Generalized Referring Expression Comprehension (GREC) and Segmentation (GRES), extend the classical visual grounding paradigm by accom…
Efficient Diffusion-Based 3D Human Pose Estimation with Hierarchical Temporal Pruning
Yuquan Bi, Hongsong Wang, Xinli Shi +3
Diffusion models have demonstrated strong capabilities in generating high-fidelity 3D human poses, yet their iterative nature and multi-hypothesis requirements incur substantial co…
Probabilistic Temporal Masked Attention for Cross-view Online Action Detection
Liping Xie, Yang Tan, Shicheng Jing +2
As a critical task in video sequence classification within computer vision, Online Action Detection (OAD) has garnered significant attention. The sensitivity of mainstream OAD mode…