works on

From the 1 of 22 linked papers with an AI index.

activity
20242026
most citedMemoryFormer: Minimize Transformer Computation by Removing Fully-Connected Layers

1 citations · 1 across the 4 of their papers we have counts for

collaborators
Showing cs.CVShow all

9 papers · 1 filter

cs.CV2025

ScaleNet: Scaling up Pretrained Neural Networks with Incremental Parameters

Zhiwei Hao, Jianyuan Guo, Li Shen +4

Recent advancements in vision transformers (ViTs) have demonstrated that larger models often achieve superior performance. However, training these models remains computationally in…

cs.CV2025

Post-Training Quantization for Diffusion Transformer via Hierarchical Timestep Grouping

Ning Ding, Jing Han, Yuchuan Tian +3

Diffusion Transformer (DiT) has now become the preferred choice for building image generation models due to its great generation capability. Unlike previous convolution-based UNet…

cs.CV2025

GPT4Image: Large Pre-trained Models Help Vision Models Learn Better on Perception Task

Ning Ding, Yehui Tang, Zhongqian Fu +3

The upsurge in pre-trained large models started by ChatGPT has swept across the entire deep learning community. Such powerful models demonstrate advanced generative ability and mul…

cs.CV2025

SAM-DiffSR: Structure-Modulated Diffusion Model for Image Super-Resolution

Chengcheng Wang, Zhiwei Hao, Yehui Tang +4

Diffusion-based super-resolution (SR) models have recently garnered significant attention due to their potent restoration capabilities. But conventional diffusion models perform no…

cs.CV2024

Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMs

Kai Han, Jianyuan Guo, Yehui Tang +3

Vision-language large models have achieved remarkable success in various multi-modal tasks, yet applying them to video understanding remains challenging due to the inherent complex…

cs.CV2024

Token Compensator: Altering Inference Cost of Vision Transformer without Re-Tuning

Shibo Jie, Yehui Tang, Jianyuan Guo +3

Token compression expedites the training and inference of Vision Transformers (ViTs) by reducing the number of the redundant tokens, e.g., pruning inattentive tokens or merging sim…