activity
20172024
most citedDynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification

311 citations · 1k across the 56 of their papers we have counts for

collaborators
Showing cs.CVShow all

64 papers · 1 filter

cs.CV2024

FlowTurbo: Towards Real-time Flow-Based Image Generation with Velocity Refiner

Wenliang Zhao, Minglei Shi, Xumin Yu +2

Building on the success of diffusion models in visual generation, flow-based models reemerge as another prominent family of generative models that have achieved competitive or bett…

cs.CV2024

DC-Solver: Improving Predictor-Corrector Diffusion Sampler via Dynamic Compensation

Wenliang Zhao, Haolin Wang, Jie Zhou +1

Diffusion probabilistic models (DPMs) have shown remarkable performance in visual synthesis but are computationally expensive due to the need for multiple evaluations during the sa…

cs.CV2022

Token-Label Alignment for Vision Transformers

Han Xiao, Wenzhao Zheng, Zheng Zhu +2

Data mixing strategies (e.g., CutMix) have shown the ability to greatly improve the performance of convolutional neural networks (CNNs). They mix two images as inputs for training…

cs.CV20221 cited

OPERA: Omni-Supervised Representation Learning with Hierarchical Supervisions

Chengkun Wang, Wenzhao Zheng, Zheng Zhu +2

The pretrain-finetune paradigm in modern computer vision facilitates the success of self-supervised learning, which tends to achieve better transferability than supervised learning…

cs.CV20221 cited

Probabilistic Deep Metric Learning for Hyperspectral Image Classification

Chengkun Wang, Wenzhao Zheng, Xian Sun +2

This paper proposes a probabilistic deep metric learning (PDML) framework for hyperspectral image classification, which aims to predict the category of each pixel for an image capt…

cs.CV2022

SemAffiNet: Semantic-Affine Transformation for Point Cloud Segmentation

Ziyi Wang, Yongming Rao, Xumin Yu +2

Conventional point cloud semantic segmentation methods usually employ an encoder-decoder architecture, where mid-level features are locally aggregated to extract geometric informat…