activity
20222024
most citedVLAB: Enhancing Video Language Pre-training by Feature Adapting and Blending

11 citations · 17 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV2024

Multi-scale Contrastive Adaptor Learning for Segmenting Anything in Underperformed Scenes

Ke Zhou, Zhongwei Qiu, Dongmei Fu

Foundational vision models, such as the Segment Anything Model (SAM), have achieved significant breakthroughs through extensive pre-training on large-scale visual datasets. Despite…

cs.CV202311 cited

VLAB: Enhancing Video Language Pre-training by Feature Adapting and Blending

Xingjian He, Sihan Chen, Fan Ma +7

Large-scale image-text contrastive pre-training models, such as CLIP, have been demonstrated to effectively learn high-quality multimodal representations. However, there is limited…

cs.CV20232 cited

PSVT: End-to-End Multi-person 3D Pose and Shape Estimation with Progressive Video Transformers

Zhongwei Qiu, Yang Qiansheng, Jian Wang +6

Existing methods of multi-person video 3D human Pose and Shape Estimation (PSE) typically adopt a two-stage strategy, which first detects human instances in each frame and then per…

cs.CV2022

IVT: An End-to-End Instance-guided Video Transformer for 3D Pose Estimation

Zhongwei Qiu, Qiansheng Yang, Jian Wang +1

Video 3D human pose estimation aims to localize the 3D coordinates of human joints from videos. Recent transformer-based approaches focus on capturing the spatiotemporal informatio…

cs.CV20224 cited

Dynamic Graph Reasoning for Multi-person 3D Pose Estimation

Zhongwei Qiu, Qiansheng Yang, Jian Wang +1

Multi-person 3D pose estimation is a challenging task because of occlusion and depth ambiguity, especially in the cases of crowd scenes. To solve these problems, most existing meth…

cs.CV2022

Learning Spatiotemporal Frequency-Transformer for Compressed Video Super-Resolution

Zhongwei Qiu, Huan Yang, Jianlong Fu +1

Compressed video super-resolution (VSR) aims to restore high-resolution frames from compressed low-resolution counterparts. Most recent VSR approaches often enhance an input frame…