31 citations · 87 across the 23 of their papers we have counts for
11 papers · 2 filters
Tracking with Human-Intent Reasoning
Jiawen Zhu, Zhi-Qi Cheng, Jun-Yan He +5
Advances in perception modeling have significantly improved the performance of object tracking. However, the current methods for specifying the target object in the initial frame a…
FMViT: A multiple-frequency mixing Vision Transformer
Wei Tan, Yifeng Geng, Xuansong Xie
The transformer model has gained widespread adoption in computer vision tasks in recent times. However, due to the quadratic time and memory complexity of self-attention, which is…
AnyText: Multilingual Visual Text Generation And Editing
Yuxiang Tuo, Wangmeng Xiang, Jun-Yan He +2
Diffusion model based Text-to-Image has achieved impressive achievements recently. Although current technology for synthesizing images is highly advanced and capable of generating…
Refined Temporal Pyramidal Compression-and-Amplification Transformer for 3D Human Pose Estimation
Hanbing Liu, Wangmeng Xiang, Jun-Yan He +4
Accurately estimating the 3D pose of humans in video sequences requires both accuracy and a well-structured architecture. With the success of transformers, we introduce the Refined…
Towards Deeply Unified Depth-aware Panoptic Segmentation with Bi-directional Guidance Learning
Junwen He, Yifan Wang, Lijun Wang +6
Depth-aware panoptic segmentation is an emerging topic in computer vision which combines semantic and geometric understanding for more robust scene interpretation. Recent works pur…
PoSynDA: Multi-Hypothesis Pose Synthesis Domain Adaptation for Robust 3D Human Pose Estimation
Hanbing Liu, Jun-Yan He, Zhi-Qi Cheng +8
Existing 3D human pose estimators face challenges in adapting to new datasets due to the lack of 2D-3D pose pairs in training sets. To overcome this issue, we propose \textit{Multi…