activity
20222024
most citedLightweight Pixel Difference Networks for Efficient Visual Representation Learning

54 citations · 68 across the 10 of their papers we have counts for

collaborators

10 papers

cs.CV20241 cited

Heterogeneous Graph Transformer for Multiple Tiny Object Tracking in RGB-T Videos

Qingyu Xu, Longguang Wang, Weidong Sheng +4

Tracking multiple tiny objects is highly challenging due to their weak appearance and limited features. Existing multi-object tracking algorithms generally focus on single-modality…

cs.CV2024

Mask-ControlNet: Higher-Quality Image Generation with An Additional Mask Prompt

Zhiqi Huang, Huixin Xiong, Haoyu Wang +2

Text-to-image generation has witnessed great progress, especially with the recent advancements in diffusion models. Since texts cannot provide detailed conditions like object appea…

cs.CV202454 cited

Lightweight Pixel Difference Networks for Efficient Visual Representation Learning

Zhuo Su, Jiehua Zhang, Longguang Wang +4

Recently, there have been tremendous efforts in developing lightweight Deep Neural Networks (DNNs) with satisfactory accuracy, which can enable the ubiquitous deployment of DNNs in…

cs.CV2023

Point Contrastive Prediction with Semantic Clustering for Self-Supervised Learning on Point Cloud Videos

Xiaoxiao Sheng, Zhiqiang Shen, Gang Xiao +3

We propose a unified point cloud video self-supervised learning framework for object-centric and scene-centric data. Previous methods commonly conduct representation learning at th…

cs.CV20232 cited

Masked Spatio-Temporal Structure Prediction for Self-supervised Learning on Point Cloud Videos

Zhiqiang Shen, Xiaoxiao Sheng, Hehe Fan +5

Recently, the community has made tremendous progress in developing effective methods for point cloud video understanding that learn from massive amounts of labeled data. However, a…

cs.CV20232 cited

PointCMP: Contrastive Mask Prediction for Self-supervised Learning on Point Cloud Videos

Zhiqiang Shen, Xiaoxiao Sheng, Longguang Wang +3

Self-supervised learning can extract representations of good quality from solely unlabeled data, which is appealing for point cloud videos due to their high labelling cost. In this…