11 citations · 17 across the 6 of their papers we have counts for
6 papers
Multi-scale Contrastive Adaptor Learning for Segmenting Anything in Underperformed Scenes
Ke Zhou, Zhongwei Qiu, Dongmei Fu
Foundational vision models, such as the Segment Anything Model (SAM), have achieved significant breakthroughs through extensive pre-training on large-scale visual datasets. Despite…
VLAB: Enhancing Video Language Pre-training by Feature Adapting and Blending
Xingjian He, Sihan Chen, Fan Ma +7
Large-scale image-text contrastive pre-training models, such as CLIP, have been demonstrated to effectively learn high-quality multimodal representations. However, there is limited…
PSVT: End-to-End Multi-person 3D Pose and Shape Estimation with Progressive Video Transformers
Zhongwei Qiu, Yang Qiansheng, Jian Wang +6
Existing methods of multi-person video 3D human Pose and Shape Estimation (PSE) typically adopt a two-stage strategy, which first detects human instances in each frame and then per…
IVT: An End-to-End Instance-guided Video Transformer for 3D Pose Estimation
Zhongwei Qiu, Qiansheng Yang, Jian Wang +1
Video 3D human pose estimation aims to localize the 3D coordinates of human joints from videos. Recent transformer-based approaches focus on capturing the spatiotemporal informatio…
Dynamic Graph Reasoning for Multi-person 3D Pose Estimation
Zhongwei Qiu, Qiansheng Yang, Jian Wang +1
Multi-person 3D pose estimation is a challenging task because of occlusion and depth ambiguity, especially in the cases of crowd scenes. To solve these problems, most existing meth…
Learning Spatiotemporal Frequency-Transformer for Compressed Video Super-Resolution
Zhongwei Qiu, Huan Yang, Jianlong Fu +1
Compressed video super-resolution (VSR) aims to restore high-resolution frames from compressed low-resolution counterparts. Most recent VSR approaches often enhance an input frame…