activity
20212026
most citedOne-for-All: Bridge the Gap Between Heterogeneous Architectures in Knowledge Distillation

31 citations · 60 across the 13 of their papers we have counts for

collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV2025

ScaleNet: Scaling up Pretrained Neural Networks with Incremental Parameters

Zhiwei Hao, Jianyuan Guo, Li Shen +4

Recent advancements in vision transformers (ViTs) have demonstrated that larger models often achieve superior performance. However, training these models remains computationally in…

cs.CV2024

ADEM-VL: Adaptive and Embedded Fusion for Efficient Vision-Language Tuning

Zhiwei Hao, Jianyuan Guo, Li Shen +3

Recent advancements in multimodal fusion have witnessed the remarkable success of vision-language (VL) models, which excel in various multimodal applications such as image captioni…

cs.CV2024★ 17 cited

GhostNetV3: Exploring the Training Strategies for Compact Models

Zhenhua Liu, Zhiwei Hao, Kai Han +2

Compact neural networks are specially designed for applications on edge devices with faster inference speed yet modest performance. However, training strategies of compact models a…

cs.CV2024★ 1 cited

Data-efficient Large Vision Models through Sequential Autoregression

Jianyuan Guo, Zhiwei Hao, Chengcheng Wang +5

Training general-purpose vision models on purely sequential visual data, eschewing linguistic inputs, has heralded a new frontier in visual understanding. These models are intended…

cs.CV2024★ 5 cited

SAM-DiffSR: Structure-Modulated Diffusion Model for Image Super-Resolution

Chengcheng Wang, Zhiwei Hao, Yehui Tang +4

Diffusion-based super-resolution (SR) models have recently garnered significant attention due to their potent restoration capabilities. But conventional diffusion models perform no…

cs.CV2023★ 31 cited

One-for-All: Bridge the Gap Between Heterogeneous Architectures in Knowledge Distillation

Zhiwei Hao, Jianyuan Guo, Kai Han +4

Knowledge distillation~(KD) has proven to be a highly effective approach for enhancing model performance through a teacher-student training scheme. However, most existing distillat…